{"id":432,"date":"2026-06-04T12:24:00","date_gmt":"2026-06-04T12:24:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/eva-bench-data\/"},"modified":"2026-06-04T18:44:46","modified_gmt":"2026-06-04T18:44:46","slug":"eva-bench-data","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/eva-bench-data\/","title":{"rendered":"EVA-Bench Knowledge 2.0: 3 Domains, 121 Instruments, 213 Eventualities"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<img decoding=\"async\" src=\"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/66d0b470cc4d59dba5c70879\/asHcI2fJBCjMUMhzxsDU6.png\" alt=\"Screenshot 2026-06-03 at 4.59.53\u202fPM\"\/><\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tIntroduction<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>Voice agent failures are sometimes extremely domain-specific. A system that flawlessly processes alphanumeric affirmation codes in flight re-booking transactions may stumble when dealing with complicated insurance policies in HR methods. Completely different domains take a look at an agent&#8217;s skill to adapt to completely different vocabulary, workflow complexities and consumer expectations. So with this launch, EVA-Bench expands from one enterprise area to 3: Airline Buyer Service Administration (CSM), Enterprise IT Service Administration (ITSM), and Healthcare HR Service Supply (HRSD). Collectively they span 213 analysis eventualities throughout 121 instruments, a roughly 4x enhance in situation protection from our unique launch. Each situation was validated for solvability in opposition to three frontier fashions (OpenAI GPT-5.4, Google Gemini 3.1 Professional, and Anthropic Claude Opus 4.6) making certain the benchmark is each difficult and truthful. All three datasets are open-source and out there for obtain:<\/p>\n<p><span class=\"hljs-keyword\">from<\/span> datasets <span class=\"hljs-keyword\">import<\/span> load_dataset<\/p>\n<p>airline = load_dataset(<span class=\"hljs-string\">&#8220;ServiceNow-AI\/eva-bench&#8221;<\/span>, <span class=\"hljs-string\">&#8220;airline&#8221;<\/span>, cut up=<span class=\"hljs-string\">&#8220;take a look at&#8221;<\/span>)<\/p>\n<p>itsm = load_dataset(<span class=\"hljs-string\">&#8220;ServiceNow-AI\/eva-bench&#8221;<\/span>, <span class=\"hljs-string\">&#8220;itsm&#8221;<\/span>, cut up=<span class=\"hljs-string\">&#8220;take a look at&#8221;<\/span>)<\/p>\n<p>hrsd = load_dataset(<span class=\"hljs-string\">&#8220;ServiceNow-AI\/eva-bench&#8221;<\/span>, <span class=\"hljs-string\">&#8220;medical&#8221;<\/span>, cut up=<span class=\"hljs-string\">&#8220;take a look at&#8221;<\/span>)<\/p>\n<p>EVA-Bench is constructed for a number of audiences. When you&#8217;re evaluating a voice agent, you&#8217;ll be able to run it in opposition to a various set of life like enterprise eventualities spanning 35+ distinct workflows. When you&#8217;re constructing your personal analysis dataset, this submit describes our end-to-end technology and validation course of in sufficient element to function a sensible reference. We stroll via how every area was designed and generated and take a deep dive into the 2 new additions. We additionally preview our upcoming multilingual extension, which widens the benchmark&#8217;s attain past English-only enterprise deployments. <\/p>\n<p>\n  <img decoding=\"async\" src=\"https:\/\/img.shields.io\/badge\/Website-blue?logo=google-chrome&amp;logoColor=white\"\/><br \/>\n  <img decoding=\"async\" src=\"https:\/\/img.shields.io\/badge\/Paper-red?logo=arxiv&amp;logoColor=white\"\/><br \/>\n  <img decoding=\"async\" src=\"https:\/\/img.shields.io\/badge\/GitHub-black?logo=github\"\/><br \/>\n  <img decoding=\"async\" src=\"https:\/\/img.shields.io\/badge\/Demo-orange?logo=play&amp;logoColor=white\"\/><br \/>\n    <img decoding=\"async\" src=\"https:\/\/img.shields.io\/badge\/Dataset-yellow?logo=huggingface&amp;logoColor=white\"\/><\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tKnowledge Design Rules<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>5 ideas guided the design of the EVA-Bench datasets throughout all three domains.<\/p>\n<p>Voice-first scope. Not each enterprise workflow belongs in a voice benchmark. We began by figuring out which duties inside every area are dealt with over the telephone in observe, then chosen the commonest flows from that subset. This saved the eventualities grounded in life like name patterns.<\/p>\n<p>Realism. Software schemas have been modeled after the sorts of APIs a manufacturing platform makes use of. State of affairs insurance policies have been drawn from precise enterprise constraints. For the Healthcare HRSD area, this meant grounding eventualities in precise US healthcare coverage and administration methods, together with NPI numbers, FMLA, and insurance coverage protection, in order that the benchmark displays the area as practitioners encounter it in actual life.<\/p>\n<p>Selection. Scaling a dataset by merely repeating equivalent duties affords restricted analysis sign. To keep away from this, we outlined particular workflows for every area and sampled throughout three situation varieties: single-intent calls, multi-intent calls with as much as 4 intents in a single dialog, and adversarial calls the place callers try to bypass troubleshooting steps, misclassify urgency, or entry data they don&#8217;t seem to be approved to view. Inside single and multi-intent eventualities, we additionally included instances the place the consumer&#8217;s aim will not be satisfiable, as a result of actual name quantity will not be all happy-path, and in our expertise fashions are likely to battle extra with unsatisfiable targets than with profitable interactions.<\/p>\n<p>Authentication. Prior work, (EVA-Bench and \u03c4-Voice), has recognized authentication as probably the most constant failure factors for voice brokers. Each area in EVA-Bench contains authentication flows, and the particular mechanisms are calibrated to the duty. For instance, OTP-based elevation seems the place a manufacturing system would truly require it, not uniformly throughout all eventualities.<\/p>\n<p>Reproducibility. With out reproducible eventualities, it&#8217;s troublesome to know whether or not a rating distinction displays a real functionality hole or an artifact of how the situation performed out. We designed the dataset so that each situation has precisely one right decision path. Consumer aim building ensures the simulator all the time has the knowledge and directions it must behave persistently, and situation technology explicitly checks for and eliminates any instances the place a number of legitimate motion sequences may obtain the identical consequence.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tState of affairs Era<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>Joint technology. Eventualities are generated utilizing SyGra, a graph-based artificial knowledge technology pipeline, with GPT-5.4 because the spine. Every situation requires three collectively constant parts that are generated collectively to forestall inconsistencies that come up when parts are produced independently:<\/p>\n<p>Consumer aim. Reproducibility requires that the consumer simulator behaves the identical approach each time a situation is run. A imprecise assertion of intent doesn&#8217;t obtain this: the simulator will make completely different judgment calls throughout runs, producing inconsistent analysis indicators. To remove this, the consumer aim is structured as a call tree that covers each scenario the simulator is more likely to encounter. The consumer aim specifies precisely which issues the consumer ought to ask for together with a negotiation sequence that specifies precisely when to push again, when to ask for alternate options, and when to simply accept. Frequent edge instances, akin to whether or not to simply accept a standby flight or an alternate airport, are dealt with with specific directions reasonably than left to the simulator to interpret. The decision situation requires proof of a accomplished motion, akin to a affirmation quantity or case ID, reasonably than a verbal dedication, so the simulator stays on the decision till the motion is definitely confirmed. The result&#8217;s a consumer that behaves like a constant, life like caller reasonably than one which improvises. <\/p>\n<p>Preliminary situation database. The backend state the agent&#8217;s instruments will question and modify through the situation. Generated collectively with the consumer aim to make sure that each entity referenced within the consumer aim, akin to reserving IDs, account particulars, and authentication credentials, exists and is constant within the database. <\/p>\n<p>Anticipated closing database state (floor fact). We derive the anticipated consequence by working the technology LLM on the agent directions, consumer aim, and preliminary situation database, producing a full motion hint. Because the LLM executes write device calls, the database is up to date incrementally, and the ensuing terminal state turns into the bottom fact that verifiers examine in opposition to throughout analysis.<\/p>\n<p>Joint technology is crucial as a result of these three parts are deeply interdependent. Impartial technology would introduce silent inconsistencies, akin to a case ID referenced within the consumer aim that doesn&#8217;t exist within the situation database, which might corrupt the analysis sign solely. To implement consistency, we run a multi-stage validation loop after every technology try and feed any failures again to the technology step, which retries till all checks move. Validation proceeds in three steps. <\/p>\n<p>A structural examine validates the situation database in opposition to a Pydantic schema, catching kind errors and lacking fields.<br \/>\nLLM-based validator checks consistency throughout the situation extra holistically: whether or not user-facing particulars within the aim match the database data, whether or not cross-references are internally legitimate, and whether or not authentication knowledge is accurately configured.<br \/>\nLLM-based hint verification move checks the complete dialog hint in opposition to coverage compliance, right motion sequencing, completion of all required terminal actions, and the absence of other write paths that may introduce non-determinism.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tAdditional Validation<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>Following SyGra technology, all eventualities went via a number of rounds of handbook evaluation. Reviewers verified that: (1) insurance policies have been utilized persistently throughout eventualities inside a website; (2) consumer targets have been particular sufficient to confess precisely one right decision; (3) anticipated closing states have been internally per each the consumer aim and the preliminary database; and (4) adversarial eventualities have been accurately specified, with a clearly identifiable coverage violation. Ambiguous or inconsistent data have been corrected or discarded.<\/p>\n<p>As a closing move, we ran three frontier fashions, OpenAI GPT-5.4, Google Gemini 3.1 Professional, and Anthropic Claude Opus 4.6, on a text-only model of every situation, bypassing the audio pipeline and offering dialog transcripts straight. For each situation on which any mannequin scored zero on job completion, we manually investigated whether or not the failure mirrored real mannequin error or a dataset subject: an ambiguous coverage, an under-specified consumer aim, a bug within the device executor, or an inconsistency between the preliminary and anticipated database states. Data with recognized dataset points have been corrected or eliminated. All chosen samples have been solvable by at the least one of many frontier fashions.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tDataset Deep-Dives<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>We created three datasets on completely different enterprise domains, every chosen to focus on a definite axis of problem for voice brokers. All three require correct transcription of structured named entities over voice (e.g., affirmation codes and worker identifiers) however differ of their major problem and variety of instruments.<\/p>\n<p>Under, we deep dive into our two new datasets: Enterprise ITSM &amp; Healthcare HRSD.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/66d0b470cc4d59dba5c70879\/qmQbPHcpbwEZUPRRx_hhi.png\" alt=\"Screenshot 2026-06-03 at 4.19.42\u202fPM\"\/><\/p>\n<p><img decoding=\"async\" src=\"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/66d0b470cc4d59dba5c70879\/7ncL-APACVJmpxbvBteLZ.png\" alt=\"Screenshot 2026-06-03 at 4.25.43\u202fPM\"\/><\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tMultilingual Assist<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>English-only analysis offers restricted perception into how a voice agent will truly carry out in one other language. Speech recognition accuracy, transcription constancy, and conversational fluency could every degrade in language-specific methods which means a high-performing voice agent in English can fail utterly when deployed in different language contexts. To provide practitioners actual perception into multilingual deployments, we&#8217;re including assist for extra languages, adapting not simply the dialog language however the analysis pipeline to every goal language and tradition:<\/p>\n<p>Names of areas referenced in eventualities<br \/>\nConsumer&#8217;s names and electronic mail addresses<br \/>\nLocalized telephone numbers<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p>English State of affairs<br \/>\nFrench State of affairs<\/p>\n<p>Utterance: &#8220;Hello, I am locked out and need assistance getting again into my account.&#8221;<br \/>\nUtterance &#8220;Bonjour, mon compte est bloqu\u00e9 et j\u2019ai besoin d\u2019aide pour y acc\u00e9der \u00e0 nouveau.&#8221;<\/p>\n<p>Places: [ &#8220;downtown&#8221;, &#8220;engineering center&#8221; ]<br \/>\nareas: [ &#8220;centre-ville&#8221;, &#8220;centre d\u2019ing\u00e9nierie&#8221; ]<\/p>\n<p>Names: {&#8220;first_name&#8221;: &#8220;Marcus&#8221;, &#8220;last_name&#8221;: &#8220;Chen&#8221;}<br \/>\nNames: {&#8220;first_name&#8221;: &#8220;\u00c9ric&#8221;, &#8220;last_name&#8221;: &#8220;Nicolas&#8221;}<\/p>\n<p>E mail: &#8220;marcus.chen@instance.com&#8221;<br \/>\nE mail: &#8220;eric.nicolas@instance.com&#8221;<\/p>\n<p>Cellphone: +1-512-555-0148<br \/>\nCellphone: +33 6 19 41 27 70<\/p>\n<\/div>\n<p>This permits the consumer simulator to offer an genuine expertise within the language of alternative. Past the dataset, we&#8217;re additionally updating our metrics and judges to construct a reliable analysis throughout languages. <\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tGet the Knowledge<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>EVA-Bench is totally open-source beneath the MIT license. The dataset, analysis framework, and leaderboard are all publicly out there. Obtain the dataset and discover particular person data on the HuggingFace dataset web page. Load any of them straight with the Hugging Face datasets library:<\/p>\n<p><span class=\"hljs-keyword\">from<\/span> datasets <span class=\"hljs-keyword\">import<\/span> load_dataset<\/p>\n<p>airline = load_dataset(<span class=\"hljs-string\">&#8220;ServiceNow-AI\/eva-bench&#8221;<\/span>, <span class=\"hljs-string\">&#8220;airline&#8221;<\/span>, cut up=<span class=\"hljs-string\">&#8220;take a look at&#8221;<\/span>)<\/p>\n<p>itsm = load_dataset(<span class=\"hljs-string\">&#8220;ServiceNow-AI\/eva-bench&#8221;<\/span>, <span class=\"hljs-string\">&#8220;itsm&#8221;<\/span>, cut up=<span class=\"hljs-string\">&#8220;take a look at&#8221;<\/span>)<\/p>\n<p>hrsd = load_dataset(<span class=\"hljs-string\">&#8220;ServiceNow-AI\/eva-bench&#8221;<\/span>, <span class=\"hljs-string\">&#8220;medical&#8221;<\/span>, cut up=<span class=\"hljs-string\">&#8220;take a look at&#8221;<\/span>)<\/p>\n<p>Every report accommodates a structured consumer aim, preliminary situation database, and floor fact anticipated closing database state \u2014 every thing wanted to run a full bot-to-bot analysis. For setup directions, code, and contributing pointers, see the GitHub repo.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tCitations<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>@misc{bogavelli2026evabenchnewendtoendframework,<br \/>\n      title={EVA-Bench: A New Finish-to-end Framework for Evaluating Voice Brokers},<br \/>\n      writer={Tara Bogavelli and Gabrielle Gauthier Melan\u00e7on and Katrina Stankiewicz and Oluwanifemi Bamgbose and Fanny Riols and Hoang H. Nguyen and Raghav Mehndiratta and Lindsay Devon Brin and Joseph Marinier and Hari Subramani and Anil Madamala and Sridhar Krishna Nemala and Srinivas Sunkara},<br \/>\n      yr={2026},<br \/>\n      eprint={2605.13841},<br \/>\n      archivePrefix={arXiv},<br \/>\n      primaryClass={cs.SD},<br \/>\n      url={https:\/\/arxiv.org\/abs\/2605.13841},<br \/>\n}<\/p>\n<p>@misc{ray2026tauvoicebenchmarkingfullduplexvoice,<br \/>\n      title={$tau$-Voice: Benchmarking Full-Duplex Voice Brokers on Actual-World Domains},<br \/>\n      writer={Soham Ray and Keshav Dhandhania and Victor Barres and Karthik Narasimhan},<br \/>\n      yr={2026},<br \/>\n      eprint={2603.13686},<br \/>\n      archivePrefix={arXiv},<br \/>\n      primaryClass={cs.SD},<br \/>\n      url={https:\/\/arxiv.org\/abs\/2603.13686},<br \/>\n}<\/p>\n<p>@misc{pradhan2025sygraunifiedgraphbasedframework,<br \/>\n      title={SyGra: A Unified Graph-Based mostly Framework for Scalable Era, High quality Tagging, and Administration of Artificial Knowledge},<br \/>\n      writer={Bidyapati Pradhan and Surajit Dasgupta and Amit Kumar Saha and Omkar Anustoop and Sriram Puttagunta and Vipul Mittal and Gopal Sarda},<br \/>\n      yr={2025},<br \/>\n      eprint={2508.15432},<br \/>\n      archivePrefix={arXiv},<br \/>\n      primaryClass={cs.AI},<br \/>\n      url={https:\/\/arxiv.org\/abs\/2508.15432},<br \/>\n}<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/huggingface.co\/blog\/ServiceNow-AI\/eva-bench-data\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Introduction Voice agent failures are sometimes extremely domain-specific. A system that flawlessly processes alphanumeric affirmation codes in flight re-booking transactions may stumble when dealing with complicated insurance policies in HR methods. Completely different domains take a look at an agent&#8217;s skill to adapt to completely different vocabulary, workflow complexities and consumer expectations. So with this [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":436,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/cdn-thumbnails.huggingface.co\/social-thumbnails\/blog\/ServiceNow-AI\/eva-bench-data.png","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[5],"tags":[160,691,690,692,56],"class_list":["post-432","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-developer-ai-open-source-ecosystem","tag-data","tag-domains","tag-evabench","tag-scenarios","tag-tools"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>EVA-Bench Knowledge 2.0: 3 Domains, 121 Instruments, 213 Eventualities - Future News 24<\/title>\n<meta name=\"description\" content=\"A Blog post by ServiceNow-AI on Hugging Face\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/eva-bench-data\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"EVA-Bench Knowledge 2.0: 3 Domains, 121 Instruments, 213 Eventualities - Future News 24\" \/>\n<meta property=\"og:description\" content=\"A Blog post by ServiceNow-AI on Hugging Face\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/eva-bench-data\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-06-04T12:24:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-06-04T18:44:46+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/cdn-thumbnails.huggingface.co\/social-thumbnails\/blog\/ServiceNow-AI\/eva-bench-data.png\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/cdn-thumbnails.huggingface.co\/social-thumbnails\/blog\/ServiceNow-AI\/eva-bench-data.png\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"10 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/eva-bench-data\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/eva-bench-data\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"EVA-Bench Knowledge 2.0: 3 Domains, 121 Instruments, 213 Eventualities\",\"datePublished\":\"2026-06-04T12:24:00+00:00\",\"dateModified\":\"2026-06-04T18:44:46+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/eva-bench-data\\\/\"},\"wordCount\":1937,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/eva-bench-data\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/cdn-thumbnails.huggingface.co\\\/social-thumbnails\\\/blog\\\/ServiceNow-AI\\\/eva-bench-data.png\",\"keywords\":[\"data\",\"Domains\",\"EVABench\",\"Scenarios\",\"Tools\"],\"articleSection\":[\"Developer AI &amp; Open-Source Ecosystem\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/eva-bench-data\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/eva-bench-data\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/eva-bench-data\\\/\",\"name\":\"EVA-Bench Knowledge 2.0: 3 Domains, 121 Instruments, 213 Eventualities - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/eva-bench-data\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/eva-bench-data\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/cdn-thumbnails.huggingface.co\\\/social-thumbnails\\\/blog\\\/ServiceNow-AI\\\/eva-bench-data.png\",\"datePublished\":\"2026-06-04T12:24:00+00:00\",\"dateModified\":\"2026-06-04T18:44:46+00:00\",\"description\":\"A Blog post by ServiceNow-AI on Hugging Face\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/eva-bench-data\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/eva-bench-data\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/eva-bench-data\\\/#primaryimage\",\"url\":\"https:\\\/\\\/cdn-thumbnails.huggingface.co\\\/social-thumbnails\\\/blog\\\/ServiceNow-AI\\\/eva-bench-data.png\",\"contentUrl\":\"https:\\\/\\\/cdn-thumbnails.huggingface.co\\\/social-thumbnails\\\/blog\\\/ServiceNow-AI\\\/eva-bench-data.png\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/eva-bench-data\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"EVA-Bench Knowledge 2.0: 3 Domains, 121 Instruments, 213 Eventualities\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"EVA-Bench Knowledge 2.0: 3 Domains, 121 Instruments, 213 Eventualities - Future News 24","description":"A Blog post by ServiceNow-AI on Hugging Face","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/eva-bench-data\/","og_locale":"en_US","og_type":"article","og_title":"EVA-Bench Knowledge 2.0: 3 Domains, 121 Instruments, 213 Eventualities - Future News 24","og_description":"A Blog post by ServiceNow-AI on Hugging Face","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/eva-bench-data\/","og_site_name":"Future News 24","article_published_time":"2026-06-04T12:24:00+00:00","article_modified_time":"2026-06-04T18:44:46+00:00","og_image":[{"url":"https:\/\/cdn-thumbnails.huggingface.co\/social-thumbnails\/blog\/ServiceNow-AI\/eva-bench-data.png","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/cdn-thumbnails.huggingface.co\/social-thumbnails\/blog\/ServiceNow-AI\/eva-bench-data.png","twitter_misc":{"Written by":"Future News 24","Est. reading time":"10 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/eva-bench-data\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/eva-bench-data\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"EVA-Bench Knowledge 2.0: 3 Domains, 121 Instruments, 213 Eventualities","datePublished":"2026-06-04T12:24:00+00:00","dateModified":"2026-06-04T18:44:46+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/eva-bench-data\/"},"wordCount":1937,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/eva-bench-data\/#primaryimage"},"thumbnailUrl":"https:\/\/cdn-thumbnails.huggingface.co\/social-thumbnails\/blog\/ServiceNow-AI\/eva-bench-data.png","keywords":["data","Domains","EVABench","Scenarios","Tools"],"articleSection":["Developer AI &amp; Open-Source Ecosystem"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/eva-bench-data\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/eva-bench-data\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/eva-bench-data\/","name":"EVA-Bench Knowledge 2.0: 3 Domains, 121 Instruments, 213 Eventualities - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/eva-bench-data\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/eva-bench-data\/#primaryimage"},"thumbnailUrl":"https:\/\/cdn-thumbnails.huggingface.co\/social-thumbnails\/blog\/ServiceNow-AI\/eva-bench-data.png","datePublished":"2026-06-04T12:24:00+00:00","dateModified":"2026-06-04T18:44:46+00:00","description":"A Blog post by ServiceNow-AI on Hugging Face","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/eva-bench-data\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/eva-bench-data\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/eva-bench-data\/#primaryimage","url":"https:\/\/cdn-thumbnails.huggingface.co\/social-thumbnails\/blog\/ServiceNow-AI\/eva-bench-data.png","contentUrl":"https:\/\/cdn-thumbnails.huggingface.co\/social-thumbnails\/blog\/ServiceNow-AI\/eva-bench-data.png"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/eva-bench-data\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"EVA-Bench Knowledge 2.0: 3 Domains, 121 Instruments, 213 Eventualities"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/432","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=432"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/432\/revisions"}],"predecessor-version":[{"id":435,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/432\/revisions\/435"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/436"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=432"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=432"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=432"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}