Wednesday, September 16, 2026
No Result
View All Result
Future News 24
Advertisement
  • Home
  • AI Research
  • Platforms
  • Ethics
  • Developer AI
  • Industry
  • Data Science
  • Emerging Tech
  • Quantum
  • BioTech
  • Decentralized
  • Home
  • AI Research
  • Platforms
  • Ethics
  • Developer AI
  • Industry
  • Data Science
  • Emerging Tech
  • Quantum
  • BioTech
  • Decentralized
No Result
View All Result
Future News 24
No Result
View All Result
Home AI Research & Breakthroughs

[2603.29139] SciVisAgentBench: A Benchmark for Evaluating Scientific Knowledge Evaluation and Visualization Brokers

Future News 24 by Future News 24
July 20, 2026
in AI Research & Breakthroughs
0 0
0
[2603.29139] SciVisAgentBench: A Benchmark for Evaluating Scientific Knowledge Evaluation and Visualization Brokers
0
SHARES
0
VIEWS
Share on FacebookShare on Twitter


[Submitted on 31 Mar 2026 (v1), last revised 16 Jul 2026 (this version, v3)]
Authors:Kuangshi Ai, Haichao Miao, Kaiyuan Tang, Nathaniel Gorski, Jianxin Solar, Guoxi Liu, Helgi I. Ingolfsson, David Lenz, Hanqi Guo, Hongfeng Yu, Teja Leburu, Michael Molash, Bei Wang, Tom Peterka, Chaoli Wang, Shusen Liu

View a PDF of the paper titled SciVisAgentBench: A Benchmark for Evaluating Scientific Knowledge Evaluation and Visualization Brokers, by Kuangshi Ai and 15 different authors

View PDF
HTML (experimental)

Summary:Current advances in massive language fashions (LLMs) have enabled agentic methods to translate natural-language intent into executable scientific visualization (SciVis) duties. Regardless of fast progress, the group lacks a principled and reproducible benchmark for evaluating these rising SciVis brokers in life like, multi-step evaluation settings. We current SciVisAgentBench, a complete and extensible benchmark for evaluating scientific knowledge evaluation and visualization brokers. Our benchmark is grounded in a structured taxonomy spanning 4 dimensions: software area, knowledge kind, complexity degree, and visualization operation. It at the moment includes 108 expert-crafted instances protecting various SciVis situations. To allow dependable evaluation, we introduce a multimodal outcome-centric analysis pipeline that mixes LLM-based judging with deterministic evaluators, together with image-based metrics, code checkers, rule-based verifiers, and case-specific evaluators. We additionally conduct a validity research with 12 SciVis specialists to look at the settlement between human and LLM judges. Utilizing this framework, we consider consultant SciVis brokers and general-purpose coding brokers to determine preliminary baselines and reveal functionality gaps. SciVisAgentBench is designed as a residing benchmark to help systematic comparability, diagnose failure modes, and drive progress in agentic SciVis. The benchmark is offered at this https URL.

Submission historical past

From: Kuangshi Ai [view email] [v1]
Tue, 31 Mar 2026 01:41:28 UTC (38,545 KB)
[v2]
Fri, 26 Jun 2026 20:48:37 UTC (40,668 KB)
[v3]
Thu, 16 Jul 2026 18:43:17 UTC (40,668 KB)



Source link

Tags: AgentsAnalysisBenchmarkdataEvaluatingscientificSciVisAgentBenchVisualization
Previous Post

AI-altered pictures on birdwatching boards placing analysis in danger | AI (synthetic intelligence)

Next Post

Full Information to Considering Machines Inkling

Next Post
Full Information to Considering Machines Inkling

Full Information to Considering Machines Inkling

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Fetching latest news…
FUTURENEWS24
Live Feed
All
AI
Dev
Industry
Frontier
Updates in 60s
FN24 AI & Tech
View All →
Future News 24

The world's leading source for AI research, emerging technology, and the people building the future. Independent, rigorous, and always ahead.

CATEGORIES

  • AI Platforms & Apps
  • AI Research & Breakthroughs
  • BioTechnology
  • Data Science & MLOps
  • Decentralized Technology
  • Developer AI & Open-Source Ecosystem
  • Emerging Technologies & Innovations
  • Ethics & Policy
  • Industry & Business
  • Quantum Computing
  • Uncategorized

LATEST

  • [2602.13312] PeroMAS: A Multi-agent System of Perovskite Materials Discovery
  • GPT-6 Astra overview: code overview good points, privateness, and value
  • GPT-6 Astra: Options, Benchmarks, Pricing, and What’s New
  • About Us
  • Advertise with Us
  • Disclaimer
  • Privacy Policy
  • DMCA 
  • Cookie Policy
  • Terms and Conditions
  • Contact us

© 2026 Future News 24. All rights reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • AI Research
  • Platforms
  • Ethics
  • Developer AI
  • Industry
  • Data Science
  • Emerging Tech
  • Quantum
  • BioTech
  • Decentralized

© 2026 Future News 24. All rights reserved.

Website security powered by MilesWeb