[Submitted on 16 May 2026]
Abstract:Peer review is an essential process in scientific research, yet the growing workload has made its automation increasingly necessary. In this study, we analyze how different types of reviewer guidelines, such as official conference guidelines and reviewer-imitating ones generated from high-quality human reviews using LLMs, affect automated peer review. Our experiments show that official conference guidelines produce review results most consistent with human judgments, suggesting that evaluation criteria refined through conference practice serve as effective guidance for automated reviewing as well. In contrast, reviewer-imitating guidelines were generally less effective than official conference guidelines. Furthermore, enforcing strict rubric-style scoring consistently degraded performance, highlighting the importance of allowing subjective and holistic scoring.
Submission history
From: Masafumi Oyamada [view email]
[v1]
Sat, 16 May 2026 02:12:55 UTC (405 KB)
0 Comments
Log in to join the conversation.No comments yet. Be the first to share your thoughts.