View a PDF of the paper titled SciVisAgentBench: A Benchmark for Evaluating Scientific Knowledge Evaluation and Visualization Brokers, by Kuangshi Ai and 15 different authors
View PDF
HTML (experimental)
Summary:Current advances in massive language fashions (LLMs) have enabled agentic methods to translate natural-language intent into executable scientific visualization (SciVis) duties. Regardless of fast progress, the group lacks a principled and reproducible benchmark for evaluating these rising SciVis brokers in life like, multi-step evaluation settings. We current SciVisAgentBench, a complete and extensible benchmark for evaluating scientific knowledge evaluation and visualization brokers. Our benchmark is grounded in a structured taxonomy spanning 4 dimensions: software area, knowledge kind, complexity degree, and visualization operation. It at the moment includes 108 expert-crafted instances protecting various SciVis situations. To allow dependable evaluation, we introduce a multimodal outcome-centric analysis pipeline that mixes LLM-based judging with deterministic evaluators, together with image-based metrics, code checkers, rule-based verifiers, and case-specific evaluators. We additionally conduct a validity research with 12 SciVis specialists to look at the settlement between human and LLM judges. Utilizing this framework, we consider consultant SciVis brokers and general-purpose coding brokers to determine preliminary baselines and reveal functionality gaps. SciVisAgentBench is designed as a residing benchmark to help systematic comparability, diagnose failure modes, and drive progress in agentic SciVis. The benchmark is offered at this https URL.
Submission historical past
From: Kuangshi Ai [view email] [v1]
Tue, 31 Mar 2026 01:41:28 UTC (38,545 KB)
[v2]
Fri, 26 Jun 2026 20:48:37 UTC (40,668 KB)
[v3]
Thu, 16 Jul 2026 18:43:17 UTC (40,668 KB)
![[2603.29139] SciVisAgentBench: A Benchmark for Evaluating Scientific Knowledge Evaluation and Visualization Brokers [2603.29139] SciVisAgentBench: A Benchmark for Evaluating Scientific Knowledge Evaluation and Visualization Brokers](http://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png)
