View a PDF of the paper titled Studying What to Keep in mind: Observability-Secure Reminiscence Retention by way of Constrained Optimization for Lengthy-Horizon Language Brokers, by Qingcan Kang and 5 different authors
View PDF
HTML (experimental)
Summary:Lengthy-horizon language brokers accumulate observations, reasoning traces, and retrieved information exceeding context home windows, making reminiscence retention a basic resource-allocation downside. Present programs deal with retention as native and don’t mannequin long-term penalties below observability constraints. To fill this hole, we formulate reminiscence retention as a constrained stochastic optimization with funds feasibility, proof utility, and delayed prices together with miss, reacquisition, and off penalties. We present this multi-step downside is NP-hard, making actual answer intractable. Furthermore, deployment choices should be made below partial observability. To handle these challenges, we suggest OSL-MR (Observability-Secure Studying for Reminiscence Retention), a learning-augmented framework that enforces a strict separation between online-observable options and offline-available supervision. OSL-MR combines an proof learner skilled from realized proof with a Blended-Rating heuristic that serves as a deployable online-safe baseline and an inductive prior. The coverage learns query-conditioned proof from interplay information and stays deployable below the identical constraints. Experiments on LoCoMo and LongMemEval present OSL-MR outperforms recency-based, Generative Brokers-style, and different heuristic baselines, particularly below tight budgets. The Blended-Rating prior improves precision and recall, and sensitivity evaluation exhibits robustness throughout price settings. On small solvable cases, single-step optimization is inadequate to anticipate future demand shifts, whereas OSL-MR stays considerably nearer to the dynamic-programming optimum, confirming the need of the sequential formulation and reinforcing our learning-guided approximation. These outcomes set up constrained stochastic optimization and optimization-guided studying as a principled basis for reminiscence administration in long-horizon brokers.
Submission historical past
From: Mingyang Liu [view email] [v1]
Tue, 9 Jun 2026 09:15:33 UTC (4,277 KB)
[v2]
Thu, 11 Jun 2026 09:47:38 UTC (4,283 KB)
[v3]
Tue, 16 Jun 2026 14:01:50 UTC (4,297 KB)
[v4]
Thu, 18 Jun 2026 02:39:41 UTC (4,299 KB)
![[2606.10616] Studying What to Keep in mind: Observability-Secure Reminiscence Retention by way of Constrained Optimization for Lengthy-Horizon Language Brokers [2606.10616] Studying What to Keep in mind: Observability-Secure Reminiscence Retention by way of Constrained Optimization for Lengthy-Horizon Language Brokers](http://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png)
