View a PDF of the paper titled FastSLM: Hierarchical Temporal Abstraction for Environment friendly Lengthy-Type Speech Adaptation, by Junseok Lee and 1 different authors
View PDF
HTML (experimental)
Summary:Scaling Multimodal Massive Language Fashions (MLLMs) to long-form speech is bottlenecked by the explosive progress of enter tokens. Present speech-language fashions undertaking high-frame-rate acoustic options immediately into the LLM enter house, making long-context processing computationally prohibitive. In contrast to photographs or movies, speech lacks spatial redundancy, making excessive token compression notably difficult. To deal with this limitation, we suggest FastSLM, a token-efficient structure that includes the Hierarchical Temporal Abstractor (HTA), which progressively distills acoustic options throughout a number of temporal scales. HTA achieves an excessive compression charge of 1.67 tokens per second (97% discount) whereas preserving important linguistic info for downstream speech-language understanding. Experimental outcomes reveal that FastSLM achieves aggressive efficiency throughout various speech-language duties whereas requiring considerably fewer speech tokens and FLOPs than present speech-language fashions. The supply code and mannequin checkpoints can be found at this https URL.
Submission historical past
From: Junseok Lee [view email] [v1]
Thu, 8 Jan 2026 07:46:03 UTC (1,898 KB)
[v2]
Mon, 2 Feb 2026 06:22:57 UTC (1,910 KB)
[v3]
Mon, 1 Jun 2026 03:39:22 UTC (1,861 KB)
[v4]
Fri, 28 Aug 2026 01:40:26 UTC (1,861 KB)
![[2601.06199] FastSLM: Hierarchical Temporal Abstraction for Environment friendly Lengthy-Type Speech Adaptation [2601.06199] FastSLM: Hierarchical Temporal Abstraction for Environment friendly Lengthy-Type Speech Adaptation](http://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png)
