Lengthy-horizon execution in Massive Language Fashions (LLMs) stays unstable even when high-level methods are offered. Evaluating on managed algorithmic puzzles, we show that whereas decomposition is crucial for stability, excessive decomposition creates a “no-recovery bottleneck”. We present that this bottleneck turns into important resulting from extremely non-uniform error distribution, the place constant errors on a number of “laborious” steps turn into irreversible. To deal with this, we suggest Lookahead-Enhanced Atomic Decomposition (LEAD). By incorporating short-horizon future validation and aggregating overlapping rollouts, LEAD gives sufficient isolation to keep up stability whereas retaining sufficient native context to right errors. This permits the o4-mini mannequin to resolve Checkers Leaping as much as complexity n = 13, whereas excessive decomposition fails past n = 11.

