{"id":1068,"date":"2026-06-12T15:56:00","date_gmt":"2026-06-12T15:56:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/06\/12\/olmo-eval\/"},"modified":"2026-06-16T09:59:26","modified_gmt":"2026-06-16T09:59:26","slug":"olmo-eval","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/06\/12\/olmo-eval\/","title":{"rendered":"olmo-eval: An analysis workbench for the mannequin growth loop"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<div class=\"not-prose\">\n<div class=\"SVELTE_HYDRATER contents\" data-target=\"BlogAuthorsByline\" data-props=\"{&quot;authors&quot;:[{&quot;author&quot;:{&quot;_id&quot;:&quot;65e5fefbe8cbae176d9ca005&quot;,&quot;avatarUrl&quot;:&quot;https:\/\/cdn-avatars.huggingface.co\/v1\/production\/uploads\/65e5fefbe8cbae176d9ca005\/5SLREDsVzycEwVsPNv765.jpeg&quot;,&quot;fullname&quot;:&quot;Tyler Murray&quot;,&quot;name&quot;:&quot;undfined&quot;,&quot;type&quot;:&quot;user&quot;,&quot;isPro&quot;:false,&quot;isHf&quot;:false,&quot;isHfAdmin&quot;:false,&quot;isMod&quot;:false,&quot;followerCount&quot;:5,&quot;isUserFollowing&quot;:false},&quot;org&quot;:{&quot;_id&quot;:&quot;5e70f3648ce3c604d78fe132&quot;,&quot;avatarUrl&quot;:&quot;https:\/\/cdn-avatars.huggingface.co\/v1\/production\/uploads\/652db071b62cf1f8463221e2\/CxxwFiaomTa1MCX_B7-pT.png&quot;,&quot;fullname&quot;:&quot;Ai2&quot;,&quot;name&quot;:&quot;allenai&quot;,&quot;type&quot;:&quot;org&quot;,&quot;isHf&quot;:false,&quot;isHfAdmin&quot;:false,&quot;isMod&quot;:false,&quot;plan&quot;:&quot;enterprise&quot;,&quot;followerCount&quot;:6085,&quot;isUserFollowing&quot;:false}},{&quot;author&quot;:{&quot;_id&quot;:&quot;638e39b249de7ae552d977b5&quot;,&quot;avatarUrl&quot;:&quot;\/avatars\/fee5cceec7536851d7c6712760716a71.svg&quot;,&quot;fullname&quot;:&quot;Kyle Wiggers&quot;,&quot;name&quot;:&quot;Ai2Comms&quot;,&quot;type&quot;:&quot;user&quot;,&quot;isPro&quot;:false,&quot;isHf&quot;:false,&quot;isHfAdmin&quot;:false,&quot;isMod&quot;:false,&quot;followerCount&quot;:13,&quot;isUserFollowing&quot;:false},&quot;org&quot;:{&quot;_id&quot;:&quot;5e70f3648ce3c604d78fe132&quot;,&quot;avatarUrl&quot;:&quot;https:\/\/cdn-avatars.huggingface.co\/v1\/production\/uploads\/652db071b62cf1f8463221e2\/CxxwFiaomTa1MCX_B7-pT.png&quot;,&quot;fullname&quot;:&quot;Ai2&quot;,&quot;name&quot;:&quot;allenai&quot;,&quot;type&quot;:&quot;org&quot;,&quot;isHf&quot;:false,&quot;isHfAdmin&quot;:false,&quot;isMod&quot;:false,&quot;plan&quot;:&quot;enterprise&quot;,&quot;followerCount&quot;:6085,&quot;isUserFollowing&quot;:false}}],&quot;translators&quot;:[],&quot;proofreaders&quot;:[],&quot;lang&quot;:&quot;en&quot;}\">\n<div class=\"not-prose\">\n<div class=\"mb-12 flex flex-wrap items-center gap-x-5 gap-y-3.5\">\n<div class=\"flex items-center font-sans leading-tight\"><span class=\"inline-block \"><span class=\"contents\"><img decoding=\"async\" class=\"rounded-full! m-0 mr-2.5 size-9 sm:mr-3 sm:size-12\" alt=\"Tyler Murray's avatar\" src=\"https:\/\/cdn-avatars.huggingface.co\/v1\/production\/uploads\/65e5fefbe8cbae176d9ca005\/5SLREDsVzycEwVsPNv765.jpeg\"\/><\/span> <\/span> <\/div>\n<div class=\"flex items-center font-sans leading-tight\"><span class=\"inline-block \"><span class=\"contents\"><img decoding=\"async\" class=\"rounded-full! m-0 mr-2.5 size-9 sm:mr-3 sm:size-12\" alt=\"Kyle Wiggers's avatar\" src=\"https:\/\/huggingface.co\/avatars\/fee5cceec7536851d7c6712760716a71.svg\"\/><\/span> <\/span> <\/div>\n<\/div><\/div>\n<\/div>\n<\/div>\n<p>\ud83d\udcbb Code: https:\/\/github.com\/allenai\/olmo-eval<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/638e39b249de7ae552d977b5\/LrVpULz4nL1G-aQ4lGdSd.png\" alt=\"Ai2 Olmo-Eval Graphic Development v3\"\/><\/p>\n<p>When you&#8217;re constructing an LLM, you consider it again and again throughout many interventions. Each adjustment to its knowledge, structure, or hyperparameters \u2014 and each step up in scale \u2014 sends you again by the identical loop: including or reconfiguring benchmarks, re-running them on every new mannequin checkpoint, noting the outcomes, and checking whether or not one thing that helped in a small experiment nonetheless holds up on the total coaching run.<\/p>\n<p>Most analysis instruments aren&#8217;t designed for this\u2014they\u2019re both constructed to run established benchmarks throughout completed fashions or run a mannequin by multi-step, tool-using issues in a sandbox. They don\u2019t sustain with a mannequin that is consistently altering, nor do they replicate how a mannequin would possibly behave below particular real-world circumstances.<\/p>\n<p>Our final venture to deal with this analysis problem was OLMES, the Open Language Mannequin Analysis Commonplace. Launched in 2024, it was meant to make LLM benchmark scores simpler to match throughout releases. The identical fashions have been being scored on the identical benchmarks in several methods \u2014 facets like immediate formatting and job formulation usually assorted from paper to paper \u2014 so claims about which fashions carried out finest usually weren&#8217;t reproducible. OLMES pinned benchmarking decisions down in an open, documented normal, and it turned the idea for evaluating our open fashions from Olmo to Tulu.<\/p>\n<p>However a mannequin&#8217;s remaining rating is simply a part of the analysis course of\u2014which is why we&#8217;re releasing olmo-eval, a brand new workbench that builds on OLMES and extends it throughout the remainder of LLM growth. In comparison with OLMES, olmo-eval cuts down the work of implementing new evaluations, affords extra flexibility in defining the place and the way they run, and makes it simpler to compose particular person parts into bigger workflows. Agentic and multi-turn analysis is supported as a first-class use case, and stronger evaluation instruments assist you decide whether or not an intervention really improved on the baseline or the distinction quantities to noise.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tHow olmo-eval differs from current instruments<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p><img decoding=\"async\" src=\"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/638e39b249de7ae552d977b5\/Um1iTvUD-lxxWYBvsNiOx.png\" alt=\"olmo-eval blog Kyles draft - Google Docs-image-1 (1)\"\/><\/p>\n<p>Is a 2.4pp change in efficiency sufficient to make a name?<\/p>\n<p>olmo-eval overlaps in some methods with Harbor, an open framework for evaluating AI brokers inside containerized, sandboxed environments. However the two instruments differ of their scope. Harbor is aimed primarily at operating and publishing agent benchmarks; olmo-eval was constructed for the on a regular basis work of growing a mannequin\u2014including and configuring benchmarks, operating them throughout checkpoints, and analyzing the outcomes immediate by immediate as a substitute of as a single total rating.<\/p>\n<p>Harbor runs the whole lot the identical means\u2014inside sealed, reproducible containers. As a result of containers will be resource-intensive, olmo-eval allows you to select how every benchmark runs as a substitute. A benchmark that simply wants a mannequin to reply questions can run instantly, which is quicker and cheaper; a benchmark that wants a locked-down setting \u2014 say, one which runs code the mannequin wrote \u2014 will get an remoted container setup. The light-weight path is the default, and olmo-eval solely opts for the heavy setup when a benchmark really requires it.<\/p>\n<p>Harbor&#8217;s course of for including a benchmark is constructed for evals you propose to publish and share publicly, with the additional verification steps that entails. olmo-eval is constructed for shifting rapidly whilst you develop, and the way you add a benchmark relies on what the benchmark wants: a brief definition for a primary eval, with choices to let a mannequin use instruments as it really works by a benchmark, or \u2014 for a benchmark that already has its personal code and process \u2014 a skinny wrapper so olmo-eval can run it as is and report the outcomes alongside different benchmark scores in the identical format.<\/p>\n<p>Each Harbor and olmo-eval preserve benchmarks separate from the runtime coverage (how the mannequin is run to provide its solutions) so you&#8217;ll be able to change one with out rewriting the opposite, however olmo-eval is designed for higher modularity. In olmo-eval, the mannequin being evaluated, the instruments it may well use, the containerized setting, and any helper fashions \u2013 like an LLM-as-a-judge \u2013 are all swappable parts. You possibly can reuse a device throughout many harnesses, or plug a grading mannequin into one benchmark with out perturbing the others, and regulate small settings (e.g., the precise wording of the immediate) with out in depth effort.<\/p>\n<p>Harbor stories an total rating for every mannequin. olmo-eval stories these scores too, every with a regular error and a minimal detectable impact (the smallest distinction that may be reliably distinguished from noise). However the extra helpful view traces the identical questions up throughout two mannequin checkpoints and compares them one after the other, with all else held fastened. This lets you see whether or not a tiny change in an total common would possibly point out an actual enchancment or just noise.<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p>Should you&#8217;re searching for&#8230;<br \/>\nolmo-eval affords<\/p>\n<p>Authoring a multi-example benchmark<br \/>\nJob subclass with a DataSource, metrics, and scoring floor<\/p>\n<p>Wrapping an current agent-style benchmark with its personal runner<br \/>\nExternalEval or SandboxedExternalEval; the benchmark retains its loop and scoring, and outcomes land in olmo-eval&#8217;s schema<\/p>\n<p>Swapping the runtime below a set benchmark<br \/>\n&#8211;harness and harness presets; the harness carries supplier, instruments, scaffold, sandboxes, and auxiliary suppliers<\/p>\n<p>Parallel container execution<br \/>\nSandbox cases for parallel executors with capability-based routing, Docker or Modal modes<\/p>\n<p>Device definitions reusable throughout duties and harnesses<br \/>\n@device decorator with non-compulsory international registry<\/p>\n<p>Multi-turn execution loops<br \/>\nScaffolds, e.g., openai_agents, chosen per harness, not baked into the duty definition<\/p>\n<\/div>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tAn built-in analysis stack<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>olmo-eval consists of 4 parts which might be helpful on their very own however designed to work collectively to tighten the experimental LLM growth loop:<\/p>\n<p>A job\/suite\/harness abstraction that decouples benchmark logic from runtime coverage. A job is the way you outline a benchmark in olmo-eval\u2014what&#8217;s being evaluated. A set teams duties right into a set you run collectively, and a harness controls how every job is run. This separation lets the identical job run as a regular baseline or with instruments and scaffolding, with out altering what it measures.<\/p>\n<p>A sandbox and capability-routing layer, together with an asynchronous sandbox planner. This helps evaluations the place a mannequin&#8217;s response relies on the actions it takes utilizing instruments, like writing and operating code or shopping the net. The purpose is to guage the mannequin&#8217;s actual device use: when a benchmark requires instruments, olmo-eval runs these instruments and feeds the outcomes again to the mannequin.<\/p>\n<p>A normalized experiment schema that data each run, its configuration, and the leads to the identical structured format. This makes it doable to group associated experiments, examine checkpoints over time, and keep away from the inconsistencies that always accumulate in long-running mannequin growth workflows.<\/p>\n<p>A outcomes viewer for pairwise mannequin comparability: lining two fashions or checkpoints up query by query surfaces small however actual efficiency modifications that an total common can cover.<\/p>\n<p>In most mannequin analysis setups, including a benchmark is a sizeable integration venture. In olmo-eval, all that\u2019s wanted is a job\u2014duties outline the benchmark dataset, how analysis requests are constructed, and the way mannequin solutions are scored (all code in Python):<\/p>\n<p><span class=\"hljs-keyword\">from<\/span> olmo_eval.widespread.formatters <span class=\"hljs-keyword\">import<\/span> ChatFormatter<br \/>\n<span class=\"hljs-keyword\">from<\/span> olmo_eval.widespread.metrics <span class=\"hljs-keyword\">import<\/span> AccuracyMetric<br \/>\n<span class=\"hljs-keyword\">from<\/span> olmo_eval.widespread.scorers <span class=\"hljs-keyword\">import<\/span> ExactMatchScorer<br \/>\n<span class=\"hljs-keyword\">from<\/span> olmo_eval.widespread.sorts <span class=\"hljs-keyword\">import<\/span> Occasion, SamplingParams<br \/>\n<span class=\"hljs-keyword\">from<\/span> olmo_eval.knowledge <span class=\"hljs-keyword\">import<\/span> DataLoader, DataSource<br \/>\n<span class=\"hljs-keyword\">from<\/span> olmo_eval.evals.duties.widespread <span class=\"hljs-keyword\">import<\/span> Job, register, register_variant<\/p>\n<p><span class=\"hljs-meta\">@register(<span class=\"hljs-params\"><span class=\"hljs-string\">&#8220;internal_freshqa&#8221;<\/span><\/span>)<\/span><br \/>\n<span class=\"hljs-keyword\">class<\/span> <span class=\"hljs-title class_\">InternalFreshQA<\/span>(<span class=\"hljs-title class_ inherited__\">Job<\/span>):<br \/>\n    data_source = DataSource(path=<span class=\"hljs-string\">&#8220;s3:\/\/evals\/inside\/freshqa.jsonl&#8221;<\/span>, break up=<span class=\"hljs-string\">&#8220;take a look at&#8221;<\/span>)<br \/>\n    formatter = ChatFormatter()<br \/>\n    sampling_params = SamplingParams(temperature=<span class=\"hljs-number\">0.0<\/span>)<br \/>\n    metrics = (AccuracyMetric(scorer=ExactMatchScorer),)<\/p>\n<p><span class=\"hljs-meta\">    @property<\/span><br \/>\n    <span class=\"hljs-keyword\">def<\/span> <span class=\"hljs-title function_\">cases<\/span>(<span class=\"hljs-params\">self<\/span>):<br \/>\n        loader = DataLoader()<br \/>\n        <span class=\"hljs-keyword\">for<\/span> idx, doc <span class=\"hljs-keyword\">in<\/span> <span class=\"hljs-built_in\">enumerate<\/span>(loader.load(self.config.get_data_source())):<br \/>\n            <span class=\"hljs-keyword\">yield<\/span> Occasion(<br \/>\n                query=doc[<span class=\"hljs-string\">&#8220;question&#8221;<\/span>],<br \/>\n                gold_answer=doc[<span class=\"hljs-string\">&#8220;answer&#8221;<\/span>],<br \/>\n                metadata={<span class=\"hljs-string\">&#8220;id&#8221;<\/span>: doc.get(<span class=\"hljs-string\">&#8220;id&#8221;<\/span>, <span class=\"hljs-string\">f&#8221;freshqa_<span class=\"hljs-subst\">{idx}<\/span>&#8220;<\/span>)},<br \/>\n            )<\/p>\n<p>Variants specific modifications in analysis coverage with out duplicating the benchmark:<\/p>\n<p>register_variant(<span class=\"hljs-string\">&#8220;internal_freshqa&#8221;<\/span>, <span class=\"hljs-string\">&#8220;3shot&#8221;<\/span>, num_fewshot=<span class=\"hljs-number\">3<\/span>, fewshot_seed=<span class=\"hljs-number\">1234<\/span>)<br \/>\nregister_variant(<span class=\"hljs-string\">&#8220;internal_freshqa&#8221;<\/span>, <span class=\"hljs-string\">&#8220;zero&#8221;<\/span>, num_fewshot=<span class=\"hljs-number\">0<\/span>)<\/p>\n<p>Suites group benchmarks into normal units you run collectively:<\/p>\n<p><span class=\"hljs-keyword\">from<\/span> olmo_eval.evals.suites <span class=\"hljs-keyword\">import<\/span> Suite, register<\/p>\n<p>register(Suite(<br \/>\n    title=<span class=\"hljs-string\">&#8220;base_qa_few_shot&#8221;<\/span>,<br \/>\n    duties=(<br \/>\n<span class=\"hljs-string\">&#8220;sciq:mc:3shot&#8221;<\/span>,<br \/>\n<span class=\"hljs-string\">&#8220;arc_challenge:mc:3shot&#8221;<\/span>,<br \/>\n<span class=\"hljs-string\">&#8220;internal_freshqa:mc:3shot&#8221;<\/span>,<br \/>\n    ),<br \/>\n))<\/p>\n<p>And since runtime coverage lives within the harness slightly than the duty definition, the identical benchmark will be simply rerun below completely different execution slightly than counting on whether or not a generated level observe merely appears believable.<\/p>\n<p><span class=\"hljs-meta prompt_\"\/><br \/>\n<span class=\"hljs-meta prompt_\"># <\/span><span class=\"language-bash\">Baseline<\/span><br \/>\nolmo-eval run -m my-instruct-checkpoint -t internal_freshqa:zero<br \/>\n<span class=\"hljs-meta prompt_\"\/><br \/>\n<span class=\"hljs-meta prompt_\"># <\/span><span class=\"language-bash\">Similar job, similar scoring, search\/device runtime enabled<\/span><br \/>\nolmo-eval run -m my-instruct-checkpoint -t internal_freshqa:zero &#8211;harness search_agent<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tReproducible analysis made open<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>Use olmo-eval when analysis is a part of ongoing mannequin growth slightly than a one-off run\u2014when it&#8217;s essential to run the identical benchmarks repeatedly throughout checkpoints below reproducible circumstances and examine interventions at each the combination and per-question degree.<\/p>\n<p>In case your recurring query is \u201cHow does this checkpoint differ from the final one, and the place precisely did it enhance or regress?\u201d, that\u2019s the workflow olmo-eval is constructed for.<\/p>\n<p>Reproducible analysis ought to preserve tempo with how fashions are constructed\u2014not solely how they&#8217;re scored as soon as they&#8217;re completed. olmo-eval carries the OLMES normal into lively mannequin growth, and we&#8217;re releasing it overtly so the group can construct on it.<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/huggingface.co\/blog\/allenai\/olmo-eval\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>\ud83d\udcbb Code: https:\/\/github.com\/allenai\/olmo-eval When you&#8217;re constructing an LLM, you consider it again and again throughout many interventions. Each adjustment to its knowledge, structure, or hyperparameters \u2014 and each step up in scale \u2014 sends you again by the identical loop: including or reconfiguring benchmarks, re-running them on every new mannequin checkpoint, noting the outcomes, and [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":1070,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/638e39b249de7ae552d977b5\/gacAOFYwPkpxu7cC4eeey.png","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[5],"tags":[761,555,1447,105,1445,1446],"class_list":["post-1068","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-developer-ai-open-source-ecosystem","tag-development","tag-evaluation","tag-loop","tag-model","tag-olmoeval","tag-workbench"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>olmo-eval: An analysis workbench for the mannequin growth loop - Future News 24<\/title>\n<meta name=\"description\" content=\"A Blog post by Ai2 on Hugging Face\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/06\/12\/olmo-eval\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"olmo-eval: An analysis workbench for the mannequin growth loop - Future News 24\" \/>\n<meta property=\"og:description\" content=\"A Blog post by Ai2 on Hugging Face\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/06\/12\/olmo-eval\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-06-12T15:56:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-06-16T09:59:26+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/638e39b249de7ae552d977b5\/gacAOFYwPkpxu7cC4eeey.png\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/638e39b249de7ae552d977b5\/gacAOFYwPkpxu7cC4eeey.png\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"8 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/12\\\/olmo-eval\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/12\\\/olmo-eval\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"olmo-eval: An analysis workbench for the mannequin growth loop\",\"datePublished\":\"2026-06-12T15:56:00+00:00\",\"dateModified\":\"2026-06-16T09:59:26+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/12\\\/olmo-eval\\\/\"},\"wordCount\":1577,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/12\\\/olmo-eval\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/cdn-uploads.huggingface.co\\\/production\\\/uploads\\\/638e39b249de7ae552d977b5\\\/gacAOFYwPkpxu7cC4eeey.png\",\"keywords\":[\"development\",\"Evaluation\",\"loop\",\"Model\",\"olmoeval\",\"workbench\"],\"articleSection\":[\"Developer AI &amp; Open-Source Ecosystem\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/12\\\/olmo-eval\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/12\\\/olmo-eval\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/12\\\/olmo-eval\\\/\",\"name\":\"olmo-eval: An analysis workbench for the mannequin growth loop - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/12\\\/olmo-eval\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/12\\\/olmo-eval\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/cdn-uploads.huggingface.co\\\/production\\\/uploads\\\/638e39b249de7ae552d977b5\\\/gacAOFYwPkpxu7cC4eeey.png\",\"datePublished\":\"2026-06-12T15:56:00+00:00\",\"dateModified\":\"2026-06-16T09:59:26+00:00\",\"description\":\"A Blog post by Ai2 on Hugging Face\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/12\\\/olmo-eval\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/12\\\/olmo-eval\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/12\\\/olmo-eval\\\/#primaryimage\",\"url\":\"https:\\\/\\\/cdn-uploads.huggingface.co\\\/production\\\/uploads\\\/638e39b249de7ae552d977b5\\\/gacAOFYwPkpxu7cC4eeey.png\",\"contentUrl\":\"https:\\\/\\\/cdn-uploads.huggingface.co\\\/production\\\/uploads\\\/638e39b249de7ae552d977b5\\\/gacAOFYwPkpxu7cC4eeey.png\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/12\\\/olmo-eval\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"olmo-eval: An analysis workbench for the mannequin growth loop\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"olmo-eval: An analysis workbench for the mannequin growth loop - Future News 24","description":"A Blog post by Ai2 on Hugging Face","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/06\/12\/olmo-eval\/","og_locale":"en_US","og_type":"article","og_title":"olmo-eval: An analysis workbench for the mannequin growth loop - Future News 24","og_description":"A Blog post by Ai2 on Hugging Face","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/06\/12\/olmo-eval\/","og_site_name":"Future News 24","article_published_time":"2026-06-12T15:56:00+00:00","article_modified_time":"2026-06-16T09:59:26+00:00","og_image":[{"url":"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/638e39b249de7ae552d977b5\/gacAOFYwPkpxu7cC4eeey.png","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/638e39b249de7ae552d977b5\/gacAOFYwPkpxu7cC4eeey.png","twitter_misc":{"Written by":"Future News 24","Est. reading time":"8 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/12\/olmo-eval\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/12\/olmo-eval\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"olmo-eval: An analysis workbench for the mannequin growth loop","datePublished":"2026-06-12T15:56:00+00:00","dateModified":"2026-06-16T09:59:26+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/12\/olmo-eval\/"},"wordCount":1577,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/12\/olmo-eval\/#primaryimage"},"thumbnailUrl":"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/638e39b249de7ae552d977b5\/gacAOFYwPkpxu7cC4eeey.png","keywords":["development","Evaluation","loop","Model","olmoeval","workbench"],"articleSection":["Developer AI &amp; Open-Source Ecosystem"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/06\/12\/olmo-eval\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/12\/olmo-eval\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/06\/12\/olmo-eval\/","name":"olmo-eval: An analysis workbench for the mannequin growth loop - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/12\/olmo-eval\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/12\/olmo-eval\/#primaryimage"},"thumbnailUrl":"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/638e39b249de7ae552d977b5\/gacAOFYwPkpxu7cC4eeey.png","datePublished":"2026-06-12T15:56:00+00:00","dateModified":"2026-06-16T09:59:26+00:00","description":"A Blog post by Ai2 on Hugging Face","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/12\/olmo-eval\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/06\/12\/olmo-eval\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/12\/olmo-eval\/#primaryimage","url":"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/638e39b249de7ae552d977b5\/gacAOFYwPkpxu7cC4eeey.png","contentUrl":"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/638e39b249de7ae552d977b5\/gacAOFYwPkpxu7cC4eeey.png"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/12\/olmo-eval\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"olmo-eval: An analysis workbench for the mannequin growth loop"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1068","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=1068"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1068\/revisions"}],"predecessor-version":[{"id":1069,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1068\/revisions\/1069"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/1070"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=1068"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=1068"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=1068"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}