View a PDF of the paper titled H+ Embedding: Harmonizing World and Token-Stage Retrieval with Context-Dependent Phrases, by Shusen Zhang and eight different authors
View PDF
HTML (experimental)
Summary:Terminology-intensive retrieval, particularly in medical settings, is determined by preserving multi-word entities, abbreviations, numerical constraints, and compositional ideas. Nevertheless, present representations lie at two extremes: single-vector retrievers typically over-compress native relevance indicators, whereas token-level late interplay retains each tokenizer subword at substantial indexing, storage, and scoring price. This mismatch raises a pure query: can context-dependent phrases present a helpful retrieval unit between world vectors and tokens? We introduce H+ Embedding, a unified multi-granularity retriever that predicts variable-length phrase partitions, preserves uncovered tokens as singletons, and applies importance-guided unit choice with weighted MaxSim interplay. Throughout 16 scientific, medical, and bilingual duties, its phrase retrieval department exceeds the worldwide retrieval department by 6.91 macro nDCG@10. It additionally practically matches Token whereas utilizing 13.7% fewer doc vectors and outperforms content-independent grouping guidelines underneath average vector budgets. Context-dependent phrase interplay due to this fact supplies an intermediate quality-cost level between world compression and token-level interplay for sensible retrieval methods.
Submission historical past
From: Junyi Hu [view email] [v1]
Wed, 29 Jul 2026 01:37:45 UTC (504 KB)
[v2]
Thu, 6 Aug 2026 07:56:44 UTC (504 KB)
[v3]
Fri, 7 Aug 2026 09:02:05 UTC (504 KB)
![[2608.00065] H+ Embedding: Harmonizing World and Token-Stage Retrieval with Context-Dependent Phrases [2608.00065] H+ Embedding: Harmonizing World and Token-Stage Retrieval with Context-Dependent Phrases](http://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png)
