{"id":3875,"date":"2026-08-06T00:00:00","date_gmt":"2026-08-06T00:00:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/08\/06\/deepambigqa-multihop-questions\/"},"modified":"2026-08-17T17:59:06","modified_gmt":"2026-08-17T17:59:06","slug":"deepambigqa-multihop-questions","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/08\/06\/deepambigqa-multihop-questions\/","title":{"rendered":"DeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Reply Completeness"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<p>Massive language fashions (LLMs) with built-in search instruments present robust promise in open-domain query answering (QA), but they typically wrestle to provide full reply set to advanced questions equivalent to \u201cWhich actor from the movie Warmth gained a minimum of one Academy Award?\u201d, which requires (1) distinguishing between a number of movies sharing the identical title and (2) reasoning throughout a big set of actors to collect and combine proof. Present QA benchmarks not often consider each challenges collectively. To deal with this, we introduce DEEPAMBIGQAGEN, an computerized knowledge technology pipeline that constructs QA duties grounded in textual content corpora and linked information graph, producing pure and verifiable questions that systematically embed title ambiguity and multi-step reasoning. Based mostly on this, we construct DEEPAMBIGQA, a dataset of three,600 questions requiring multi-hop reasoning and half of them express title ambiguity resolving. Experiments reveal that, even state-of-the-art GPT-5 present incomplete solutions, reaching solely 0.13 actual match on ambiguous questions and 0.21 on non-ambiguous questions. These findings spotlight the necessity for extra strong QA techniques geared toward data gathering and reply completeness.<\/p>\n<p>\u2020 College of California, Santa Barbara** Work executed whereas at Apple<\/p><\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/machinelearning.apple.com\/research\/deepambigqa-multihop-questions\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Massive language fashions (LLMs) with built-in search instruments present robust promise in open-domain query answering (QA), but they typically wrestle to provide full reply set to advanced questions equivalent to \u201cWhich actor from the movie Warmth gained a minimum of one Academy Award?\u201d, which requires (1) distinguishing between a number of movies sharing the identical [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":3877,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[2],"tags":[4194,420,474,4196,4193,452,4195,74],"class_list":["post-3875","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-research-breakthroughs","tag-ambiguous","tag-answer","tag-benchmarking","tag-completeness","tag-deepambigqa","tag-llm","tag-multihop","tag-questions"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>DeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Reply Completeness - Future News 24<\/title>\n<meta name=\"description\" content=\"Large language models (LLMs) with integrated search tools show strong promise in open-domain question answering (QA), yet they often\u2026\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/08\/06\/deepambigqa-multihop-questions\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"DeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Reply Completeness - Future News 24\" \/>\n<meta property=\"og:description\" content=\"Large language models (LLMs) with integrated search tools show strong promise in open-domain question answering (QA), yet they often\u2026\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/08\/06\/deepambigqa-multihop-questions\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-08-06T00:00:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-08-17T17:59:06+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"1 minute\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/06\\\/deepambigqa-multihop-questions\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/06\\\/deepambigqa-multihop-questions\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"DeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Reply Completeness\",\"datePublished\":\"2026-08-06T00:00:00+00:00\",\"dateModified\":\"2026-08-17T17:59:06+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/06\\\/deepambigqa-multihop-questions\\\/\"},\"wordCount\":198,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/06\\\/deepambigqa-multihop-questions\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/mlr.cdn-apple.com\\\/media\\\/Home_1200x630_48225d82e9.png\",\"keywords\":[\"Ambiguous\",\"Answer\",\"Benchmarking\",\"Completeness\",\"DeepAmbigQA\",\"LLM\",\"Multihop\",\"Questions\"],\"articleSection\":[\"AI Research &amp; Breakthroughs\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/06\\\/deepambigqa-multihop-questions\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/06\\\/deepambigqa-multihop-questions\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/06\\\/deepambigqa-multihop-questions\\\/\",\"name\":\"DeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Reply Completeness - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/06\\\/deepambigqa-multihop-questions\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/06\\\/deepambigqa-multihop-questions\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/mlr.cdn-apple.com\\\/media\\\/Home_1200x630_48225d82e9.png\",\"datePublished\":\"2026-08-06T00:00:00+00:00\",\"dateModified\":\"2026-08-17T17:59:06+00:00\",\"description\":\"Large language models (LLMs) with integrated search tools show strong promise in open-domain question answering (QA), yet they often\u2026\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/06\\\/deepambigqa-multihop-questions\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/06\\\/deepambigqa-multihop-questions\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/06\\\/deepambigqa-multihop-questions\\\/#primaryimage\",\"url\":\"https:\\\/\\\/mlr.cdn-apple.com\\\/media\\\/Home_1200x630_48225d82e9.png\",\"contentUrl\":\"https:\\\/\\\/mlr.cdn-apple.com\\\/media\\\/Home_1200x630_48225d82e9.png\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/06\\\/deepambigqa-multihop-questions\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"DeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Reply Completeness\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"DeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Reply Completeness - Future News 24","description":"Large language models (LLMs) with integrated search tools show strong promise in open-domain question answering (QA), yet they often\u2026","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/08\/06\/deepambigqa-multihop-questions\/","og_locale":"en_US","og_type":"article","og_title":"DeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Reply Completeness - Future News 24","og_description":"Large language models (LLMs) with integrated search tools show strong promise in open-domain question answering (QA), yet they often\u2026","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/08\/06\/deepambigqa-multihop-questions\/","og_site_name":"Future News 24","article_published_time":"2026-08-06T00:00:00+00:00","article_modified_time":"2026-08-17T17:59:06+00:00","og_image":[{"url":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","twitter_misc":{"Written by":"Future News 24","Est. reading time":"1 minute"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/06\/deepambigqa-multihop-questions\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/06\/deepambigqa-multihop-questions\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"DeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Reply Completeness","datePublished":"2026-08-06T00:00:00+00:00","dateModified":"2026-08-17T17:59:06+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/06\/deepambigqa-multihop-questions\/"},"wordCount":198,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/06\/deepambigqa-multihop-questions\/#primaryimage"},"thumbnailUrl":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","keywords":["Ambiguous","Answer","Benchmarking","Completeness","DeepAmbigQA","LLM","Multihop","Questions"],"articleSection":["AI Research &amp; Breakthroughs"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/08\/06\/deepambigqa-multihop-questions\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/06\/deepambigqa-multihop-questions\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/08\/06\/deepambigqa-multihop-questions\/","name":"DeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Reply Completeness - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/06\/deepambigqa-multihop-questions\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/06\/deepambigqa-multihop-questions\/#primaryimage"},"thumbnailUrl":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","datePublished":"2026-08-06T00:00:00+00:00","dateModified":"2026-08-17T17:59:06+00:00","description":"Large language models (LLMs) with integrated search tools show strong promise in open-domain question answering (QA), yet they often\u2026","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/06\/deepambigqa-multihop-questions\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/08\/06\/deepambigqa-multihop-questions\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/06\/deepambigqa-multihop-questions\/#primaryimage","url":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","contentUrl":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/06\/deepambigqa-multihop-questions\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"DeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Reply Completeness"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3875","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=3875"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3875\/revisions"}],"predecessor-version":[{"id":3876,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3875\/revisions\/3876"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/3877"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=3875"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=3875"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=3875"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}