Wednesday, September 16, 2026
No Result
View All Result
Future News 24
Advertisement
  • Home
  • AI Research
  • Platforms
  • Ethics
  • Developer AI
  • Industry
  • Data Science
  • Emerging Tech
  • Quantum
  • BioTech
  • Decentralized
  • Home
  • AI Research
  • Platforms
  • Ethics
  • Developer AI
  • Industry
  • Data Science
  • Emerging Tech
  • Quantum
  • BioTech
  • Decentralized
No Result
View All Result
Future News 24
No Result
View All Result
Home AI Research & Breakthroughs

[2602.11354] ReplicatorBench: Benchmarking LLM Brokers for Replicability in Social and Behavioral Sciences

Future News 24 by Future News 24
July 2, 2026
in AI Research & Breakthroughs
0 0
0
[2602.11354] ReplicatorBench: Benchmarking LLM Brokers for Replicability in Social and Behavioral Sciences
0
SHARES
0
VIEWS
Share on FacebookShare on Twitter


[Submitted on 11 Feb 2026 (v1), last revised 29 Jun 2026 (this version, v3)]
Authors:Bang Nguyen, Dominik Soós, Qian Ma, Rochana R. Obadage, Zack Ranjan, Sai Koneru, Timothy M. Errington, Shakhlo Nematova, Sarah Rajtmajer, Jian Wu, Meng Jiang

View a PDF of the paper titled ReplicatorBench: Benchmarking LLM Brokers for Replicability in Social and Behavioral Sciences, by Bang Nguyen and 10 different authors

View PDF
HTML (experimental)

Summary:The literature has witnessed an rising curiosity in AI brokers for automated evaluation of scientific papers. Present benchmarks focus totally on the computational side of this process, testing brokers’ means to breed or replicate analysis outcomes when accessing the code and knowledge. This setting, whereas foundational, (1) fails to seize the inconsistent availability of latest knowledge for replication versus copy, and (2) lacks ground-truth range by focusing solely on reproducible papers, thereby failing to guage an agent’s means to determine non-replicable analysis. Moreover, most benchmarks solely consider outcomes relatively than the replication course of. In response, we introduce ReplicatorBench, an end-to-end benchmark, together with human-verified replicable and non-replicable analysis claims in social and behavioral sciences for evaluating AI brokers in analysis replication throughout three levels: (1) extraction and retrieval of replication knowledge; (2) design and execution of computational experiments; and (3) interpretation of outcomes, permitting a take a look at of AI brokers’ functionality to imitate the actions of human replicators in actual world. To set a baseline of AI brokers’ functionality, we develop ReplicatorAgent, an agentic framework outfitted with essential instruments, like net search and iterative interplay with sandboxed environments, to perform duties in ReplicatorBench. We consider ReplicatorAgent throughout 4 underlying massive language fashions (LLMs), in addition to totally different design decisions of programming language and ranges of code entry. Our findings reveal that whereas present LLM brokers are able to successfully designing and executing computational experiments, they wrestle with retrieving sources, akin to new knowledge, essential to copy a declare. All code and knowledge are publicly accessible at this https URL.

Submission historical past

From: Bang Nguyen [view email] [v1]
Wed, 11 Feb 2026 20:42:10 UTC (1,064 KB)
[v2]
Thu, 9 Apr 2026 18:26:34 UTC (282 KB)
[v3]
Mon, 29 Jun 2026 19:08:09 UTC (260 KB)



Source link

Tags: AgentsBehavioralBenchmarkingLLMReplicabilityReplicatorBenchSciencesSocial
Previous Post

Ethereum for Governments and Establishments: Why impartial infrastructure issues now

Next Post

July 2026 Publication – by Piers Stobbs

Next Post
July 2026 Publication – by Piers Stobbs

July 2026 Publication - by Piers Stobbs

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Fetching latest news…
FUTURENEWS24
Live Feed
All
AI
Dev
Industry
Frontier
Updates in 60s
FN24 AI & Tech
View All →
Future News 24

The world's leading source for AI research, emerging technology, and the people building the future. Independent, rigorous, and always ahead.

CATEGORIES

  • AI Platforms & Apps
  • AI Research & Breakthroughs
  • BioTechnology
  • Data Science & MLOps
  • Decentralized Technology
  • Developer AI & Open-Source Ecosystem
  • Emerging Technologies & Innovations
  • Ethics & Policy
  • Industry & Business
  • Quantum Computing
  • Uncategorized

LATEST

  • [2602.13312] PeroMAS: A Multi-agent System of Perovskite Materials Discovery
  • GPT-6 Astra overview: code overview good points, privateness, and value
  • GPT-6 Astra: Options, Benchmarks, Pricing, and What’s New
  • About Us
  • Advertise with Us
  • Disclaimer
  • Privacy Policy
  • DMCA 
  • Cookie Policy
  • Terms and Conditions
  • Contact us

© 2026 Future News 24. All rights reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • AI Research
  • Platforms
  • Ethics
  • Developer AI
  • Industry
  • Data Science
  • Emerging Tech
  • Quantum
  • BioTech
  • Decentralized

© 2026 Future News 24. All rights reserved.

Website security powered by MilesWeb