View a PDF of the paper titled ReplicatorBench: Benchmarking LLM Brokers for Replicability in Social and Behavioral Sciences, by Bang Nguyen and 10 different authors
View PDF
HTML (experimental)
Summary:The literature has witnessed an rising curiosity in AI brokers for automated evaluation of scientific papers. Present benchmarks focus totally on the computational side of this process, testing brokers’ means to breed or replicate analysis outcomes when accessing the code and knowledge. This setting, whereas foundational, (1) fails to seize the inconsistent availability of latest knowledge for replication versus copy, and (2) lacks ground-truth range by focusing solely on reproducible papers, thereby failing to guage an agent’s means to determine non-replicable analysis. Moreover, most benchmarks solely consider outcomes relatively than the replication course of. In response, we introduce ReplicatorBench, an end-to-end benchmark, together with human-verified replicable and non-replicable analysis claims in social and behavioral sciences for evaluating AI brokers in analysis replication throughout three levels: (1) extraction and retrieval of replication knowledge; (2) design and execution of computational experiments; and (3) interpretation of outcomes, permitting a take a look at of AI brokers’ functionality to imitate the actions of human replicators in actual world. To set a baseline of AI brokers’ functionality, we develop ReplicatorAgent, an agentic framework outfitted with essential instruments, like net search and iterative interplay with sandboxed environments, to perform duties in ReplicatorBench. We consider ReplicatorAgent throughout 4 underlying massive language fashions (LLMs), in addition to totally different design decisions of programming language and ranges of code entry. Our findings reveal that whereas present LLM brokers are able to successfully designing and executing computational experiments, they wrestle with retrieving sources, akin to new knowledge, essential to copy a declare. All code and knowledge are publicly accessible at this https URL.
Submission historical past
From: Bang Nguyen [view email] [v1]
Wed, 11 Feb 2026 20:42:10 UTC (1,064 KB)
[v2]
Thu, 9 Apr 2026 18:26:34 UTC (282 KB)
[v3]
Mon, 29 Jun 2026 19:08:09 UTC (260 KB)
![[2602.11354] ReplicatorBench: Benchmarking LLM Brokers for Replicability in Social and Behavioral Sciences [2602.11354] ReplicatorBench: Benchmarking LLM Brokers for Replicability in Social and Behavioral Sciences](http://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png)
