View a PDF of the paper titled Tuning Language Fashions by Combination-of-Depths Ensemble, by Haoyan Luo and 1 different authors
View PDF
Summary:Transformer-based Massive Language Fashions (LLMs) historically depend on final-layer loss for finetuning and final-layer representations for predictions, probably overlooking the predictive energy embedded in late layers. Interpretability instruments such because the logit lens present that late-layer representations already carry largely fashioned, task-relevant predictions; right here we ask whether or not that remark could be changed into an actionable coaching sign. We discover that focusing tuning effort on these layers can yield losses akin to these of the ultimate layer, with complementary test-time behaviour. Constructing on this, we introduce a tuning framework, Combination-of-Depths Ensemble (MoDE), which treats the late layers as an ensemble that contributes to the ultimate logits by discovered routing weights. MoDE could be utilized on high of any current tuning methodology (e.g., LoRA) and, in our experiments, modestly improves reasoning efficiency at a small parameter overhead. We current MoDE as a mechanism research displaying that late-layer logits could be made instantly helpful for tuning, and that they’ll substitute for considerably bigger trainable modules with comparable efficiency.
Submission historical past
From: Haoyan Luo [view email] [v1]
Wed, 16 Oct 2024 22:51:45 UTC (12,478 KB)
[v2]
Wed, 24 Jun 2026 22:05:06 UTC (6,261 KB)
![[2410.13077] Tuning Language Fashions by Combination-of-Depths Ensemble [2410.13077] Tuning Language Fashions by Combination-of-Depths Ensemble](http://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png)
