TY - GEN
T1 - All Accept, No Reject
T2 - 2026 CHI Conference on Human Factors in Computing Systems, CHI 2026
AU - Verma, Nitin
AU - Landrum, Asheley R.
N1 - Publisher Copyright:
© 2026 Copyright held by the owner/author(s).
PY - 2026/4/13
Y1 - 2026/4/13
N2 - An exponential rise in manuscript submission volume has strained the peer review system, prompting interest in automation from overburdened scholars and publishers. We systematically evaluated GPT-4.1, GPT-4o, o1, o3, o3-mini, and GPT-5 as "peer"reviewers, comparing their evaluation and acceptance of 137 manuscripts from an open dataset (PeerRead) with the corresponding human-generated reviews. While o3's and GPT-5's acceptance rates were close to the human benchmark (∼67% of submissions), others approved nearly every paper (>98%); all models performed extremely poorly on accuracy, precision, and recall metrics. To probe this striking "yes-bias", we profiled the LLMs using Schwartz's Portrait Values Questionnaire (PVQ-RR) and found that all LLMs emphasized self-transcendence and openness-to-change and de-emphasized conservation and self-enhancement. We argue that value orientations of LLMs we investigated are misaligned with the values underpinning peer review, and suggest new research on aligning AI judgment systems with human goals in this context.
AB - An exponential rise in manuscript submission volume has strained the peer review system, prompting interest in automation from overburdened scholars and publishers. We systematically evaluated GPT-4.1, GPT-4o, o1, o3, o3-mini, and GPT-5 as "peer"reviewers, comparing their evaluation and acceptance of 137 manuscripts from an open dataset (PeerRead) with the corresponding human-generated reviews. While o3's and GPT-5's acceptance rates were close to the human benchmark (∼67% of submissions), others approved nearly every paper (>98%); all models performed extremely poorly on accuracy, precision, and recall metrics. To probe this striking "yes-bias", we profiled the LLMs using Schwartz's Portrait Values Questionnaire (PVQ-RR) and found that all LLMs emphasized self-transcendence and openness-to-change and de-emphasized conservation and self-enhancement. We argue that value orientations of LLMs we investigated are misaligned with the values underpinning peer review, and suggest new research on aligning AI judgment systems with human goals in this context.
KW - Human Values
KW - Large Language Models
KW - Peer Review
KW - Science
UR - https://www.scopus.com/pages/publications/105038728916
UR - https://www.scopus.com/pages/publications/105038728916#tab=citedBy
U2 - 10.1145/3772318.3791300
DO - 10.1145/3772318.3791300
M3 - Conference contribution
AN - SCOPUS:105038728916
T3 - Conference on Human Factors in Computing Systems - Proceedings
BT - CHI 2026 - Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems
A2 - Oliver, Nuria
A2 - Shamma, David A.
A2 - Candello, Heloisa
A2 - Cesar, Pablo
A2 - Lopes, Pedro
A2 - Bozzon, Alessandro
A2 - Kosch, Thomas
A2 - Liao, Vera
A2 - Ma, Xiaojuan
A2 - Artizzu, Valentino
A2 - Draxler, Fiona
A2 - Lopez, Gustavo
A2 - Reinschluessel, Anke V.
A2 - Tong, Xin
A2 - Toups Dugas, Phoebe O.
PB - Association for Computing Machinery
Y2 - 13 April 2026 through 17 April 2026
ER -