{"id":3545,"date":"2026-08-10T04:00:00","date_gmt":"2026-08-10T04:00:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/08\/10\/2608-00065\/"},"modified":"2026-08-10T08:59:04","modified_gmt":"2026-08-10T08:59:04","slug":"2608-00065","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/08\/10\/2608-00065\/","title":{"rendered":"[2608.00065] H+ Embedding: Harmonizing World and Token-Stage Retrieval with Context-Dependent Phrases"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div id=\"content-inner\">\n<div id=\"abs\">\n<div class=\"dateline\">\n  [Submitted on 29 Jul 2026 (v1), last revised 7 Aug 2026 (this version, v3)]<\/div>\n<p>View a PDF of the paper titled H+ Embedding: Harmonizing World and Token-Stage Retrieval with Context-Dependent Phrases, by Shusen Zhang and eight different authors<\/p>\n<p>    View PDF<br \/>\n    HTML (experimental)<\/p>\n<blockquote class=\"abstract mathjax\"><p>\n            <span class=\"descriptor\">Summary:<\/span>Terminology-intensive retrieval, particularly in medical settings, is determined by preserving multi-word entities, abbreviations, numerical constraints, and compositional ideas. Nevertheless, present representations lie at two extremes: single-vector retrievers typically over-compress native relevance indicators, whereas token-level late interplay retains each tokenizer subword at substantial indexing, storage, and scoring price. This mismatch raises a pure query: can context-dependent phrases present a helpful retrieval unit between world vectors and tokens? We introduce H+ Embedding, a unified multi-granularity retriever that predicts variable-length phrase partitions, preserves uncovered tokens as singletons, and applies importance-guided unit choice with weighted MaxSim interplay. Throughout 16 scientific, medical, and bilingual duties, its phrase retrieval department exceeds the worldwide retrieval department by 6.91 macro nDCG@10. It additionally practically matches Token whereas utilizing 13.7% fewer doc vectors and outperforms content-independent grouping guidelines underneath average vector budgets. Context-dependent phrase interplay due to this fact supplies an intermediate quality-cost level between world compression and token-level interplay for sensible retrieval methods.\n    <\/p><\/blockquote><\/div>\n<\/div>\n<div>\n<h2>Submission historical past<\/h2>\n<p> From: Junyi Hu [view email]                  [v1]<br \/>\n        Wed, 29 Jul 2026 01:37:45 UTC (504 KB)<br \/>\n            [v2]<br \/>\n        Thu, 6 Aug 2026 07:56:44 UTC (504 KB)<br \/>\n    [v3]<br \/>\n        Fri, 7 Aug 2026 09:02:05 UTC (504 KB)\n<\/p><\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/arxiv.org\/abs\/2608.00065\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>[Submitted on 29 Jul 2026 (v1), last revised 7 Aug 2026 (this version, v3)] View a PDF of the paper titled H+ Embedding: Harmonizing World and Token-Stage Retrieval with Context-Dependent Phrases, by Shusen Zhang and eight different authors View PDF HTML (experimental) Summary:Terminology-intensive retrieval, particularly in medical settings, is determined by preserving multi-word entities, abbreviations, [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":3547,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"http:\/\/arxiv.org\/static\/browse\/0.3.4\/images\/arxiv-logo-fb.png","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[2],"tags":[3910,3908,705,3909,3911,1235,3115],"class_list":["post-3545","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-research-breakthroughs","tag-contextdependent","tag-embedding","tag-global","tag-harmonizing","tag-phrases","tag-retrieval","tag-tokenlevel"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>[2608.00065] H+ Embedding: Harmonizing World and Token-Stage Retrieval with Context-Dependent Phrases - Future News 24<\/title>\n<meta name=\"description\" content=\"Abstract page for arXiv paper 2608.00065: H+ Embedding: Harmonizing Global and Token-Level Retrieval with Context-Dependent Phrases\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/08\/10\/2608-00065\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"[2608.00065] H+ Embedding: Harmonizing World and Token-Stage Retrieval with Context-Dependent Phrases - Future News 24\" \/>\n<meta property=\"og:description\" content=\"Abstract page for arXiv paper 2608.00065: H+ Embedding: Harmonizing Global and Token-Level Retrieval with Context-Dependent Phrases\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/08\/10\/2608-00065\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-08-10T04:00:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-08-10T08:59:04+00:00\" \/>\n<meta property=\"og:image\" content=\"http:\/\/arxiv.org\/static\/browse\/0.3.4\/images\/arxiv-logo-fb.png\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"http:\/\/arxiv.org\/static\/browse\/0.3.4\/images\/arxiv-logo-fb.png\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"1 minute\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/10\\\/2608-00065\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/10\\\/2608-00065\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"[2608.00065] H+ Embedding: Harmonizing World and Token-Stage Retrieval with Context-Dependent Phrases\",\"datePublished\":\"2026-08-10T04:00:00+00:00\",\"dateModified\":\"2026-08-10T08:59:04+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/10\\\/2608-00065\\\/\"},\"wordCount\":228,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/10\\\/2608-00065\\\/#primaryimage\"},\"thumbnailUrl\":\"http:\\\/\\\/arxiv.org\\\/static\\\/browse\\\/0.3.4\\\/images\\\/arxiv-logo-fb.png\",\"keywords\":[\"ContextDependent\",\"Embedding\",\"global\",\"Harmonizing\",\"Phrases\",\"retrieval\",\"TokenLevel\"],\"articleSection\":[\"AI Research &amp; Breakthroughs\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/10\\\/2608-00065\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/10\\\/2608-00065\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/10\\\/2608-00065\\\/\",\"name\":\"[2608.00065] H+ Embedding: Harmonizing World and Token-Stage Retrieval with Context-Dependent Phrases - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/10\\\/2608-00065\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/10\\\/2608-00065\\\/#primaryimage\"},\"thumbnailUrl\":\"http:\\\/\\\/arxiv.org\\\/static\\\/browse\\\/0.3.4\\\/images\\\/arxiv-logo-fb.png\",\"datePublished\":\"2026-08-10T04:00:00+00:00\",\"dateModified\":\"2026-08-10T08:59:04+00:00\",\"description\":\"Abstract page for arXiv paper 2608.00065: H+ Embedding: Harmonizing Global and Token-Level Retrieval with Context-Dependent Phrases\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/10\\\/2608-00065\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/10\\\/2608-00065\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/10\\\/2608-00065\\\/#primaryimage\",\"url\":\"http:\\\/\\\/arxiv.org\\\/static\\\/browse\\\/0.3.4\\\/images\\\/arxiv-logo-fb.png\",\"contentUrl\":\"http:\\\/\\\/arxiv.org\\\/static\\\/browse\\\/0.3.4\\\/images\\\/arxiv-logo-fb.png\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/10\\\/2608-00065\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"[2608.00065] H+ Embedding: Harmonizing World and Token-Stage Retrieval with Context-Dependent Phrases\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"[2608.00065] H+ Embedding: Harmonizing World and Token-Stage Retrieval with Context-Dependent Phrases - Future News 24","description":"Abstract page for arXiv paper 2608.00065: H+ Embedding: Harmonizing Global and Token-Level Retrieval with Context-Dependent Phrases","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/08\/10\/2608-00065\/","og_locale":"en_US","og_type":"article","og_title":"[2608.00065] H+ Embedding: Harmonizing World and Token-Stage Retrieval with Context-Dependent Phrases - Future News 24","og_description":"Abstract page for arXiv paper 2608.00065: H+ Embedding: Harmonizing Global and Token-Level Retrieval with Context-Dependent Phrases","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/08\/10\/2608-00065\/","og_site_name":"Future News 24","article_published_time":"2026-08-10T04:00:00+00:00","article_modified_time":"2026-08-10T08:59:04+00:00","og_image":[{"url":"http:\/\/arxiv.org\/static\/browse\/0.3.4\/images\/arxiv-logo-fb.png","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"http:\/\/arxiv.org\/static\/browse\/0.3.4\/images\/arxiv-logo-fb.png","twitter_misc":{"Written by":"Future News 24","Est. reading time":"1 minute"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/10\/2608-00065\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/10\/2608-00065\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"[2608.00065] H+ Embedding: Harmonizing World and Token-Stage Retrieval with Context-Dependent Phrases","datePublished":"2026-08-10T04:00:00+00:00","dateModified":"2026-08-10T08:59:04+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/10\/2608-00065\/"},"wordCount":228,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/10\/2608-00065\/#primaryimage"},"thumbnailUrl":"http:\/\/arxiv.org\/static\/browse\/0.3.4\/images\/arxiv-logo-fb.png","keywords":["ContextDependent","Embedding","global","Harmonizing","Phrases","retrieval","TokenLevel"],"articleSection":["AI Research &amp; Breakthroughs"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/08\/10\/2608-00065\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/10\/2608-00065\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/08\/10\/2608-00065\/","name":"[2608.00065] H+ Embedding: Harmonizing World and Token-Stage Retrieval with Context-Dependent Phrases - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/10\/2608-00065\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/10\/2608-00065\/#primaryimage"},"thumbnailUrl":"http:\/\/arxiv.org\/static\/browse\/0.3.4\/images\/arxiv-logo-fb.png","datePublished":"2026-08-10T04:00:00+00:00","dateModified":"2026-08-10T08:59:04+00:00","description":"Abstract page for arXiv paper 2608.00065: H+ Embedding: Harmonizing Global and Token-Level Retrieval with Context-Dependent Phrases","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/10\/2608-00065\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/08\/10\/2608-00065\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/10\/2608-00065\/#primaryimage","url":"http:\/\/arxiv.org\/static\/browse\/0.3.4\/images\/arxiv-logo-fb.png","contentUrl":"http:\/\/arxiv.org\/static\/browse\/0.3.4\/images\/arxiv-logo-fb.png"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/10\/2608-00065\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"[2608.00065] H+ Embedding: Harmonizing World and Token-Stage Retrieval with Context-Dependent Phrases"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3545","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=3545"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3545\/revisions"}],"predecessor-version":[{"id":3546,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3545\/revisions\/3546"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/3547"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=3545"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=3545"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=3545"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}