Massive language fashions (LLMs) with built-in search instruments present robust promise in open-domain query answering (QA), but they typically wrestle to provide full reply set to advanced questions equivalent to “Which actor from the movie Warmth gained a minimum of one Academy Award?”, which requires (1) distinguishing between a number of movies sharing the identical title and (2) reasoning throughout a big set of actors to collect and combine proof. Present QA benchmarks not often consider each challenges collectively. To deal with this, we introduce DEEPAMBIGQAGEN, an computerized knowledge technology pipeline that constructs QA duties grounded in textual content corpora and linked information graph, producing pure and verifiable questions that systematically embed title ambiguity and multi-step reasoning. Based mostly on this, we construct DEEPAMBIGQA, a dataset of three,600 questions requiring multi-hop reasoning and half of them express title ambiguity resolving. Experiments reveal that, even state-of-the-art GPT-5 present incomplete solutions, reaching solely 0.13 actual match on ambiguous questions and 0.21 on non-ambiguous questions. These findings spotlight the necessity for extra strong QA techniques geared toward data gathering and reply completeness.
† College of California, Santa Barbara** Work executed whereas at Apple

![[2510.21084] MediRec: Enhancing Chinese language Treatment Suggestion with Explainable Medical Reasoning [2510.21084] MediRec: Enhancing Chinese language Treatment Suggestion with Explainable Medical Reasoning](http://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png)