Publications

When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction

NeurIPS 2026

Publication date: December 12, 2026

Vardhan_Dongre, Joseph Hsieh, Viet Lai, David Seunghyun Yoon, Trung Bui, Dilek Hakkani-Tür

Large language models reliably follow complex instructions in a single turn, yet across long multi-turn interactions they start strong then gradually lose the thread of the instructions, persona, and rules they were given. This degradation has been measured behaviorally but not mechanistically explained. We trace this failure to a transition between two information channels: accessibility to goal-defining tokens through attention, and a residual channel that carries goal information forward through hidden states. We introduce the Goal Accessibility Ratio (GAR), measuring attention from generated tokens to task-defining goal tokens, and combine it with sliding-window ablations and residual-stream probes. When attention to instructions closes, what survives reveals architecture. Across the architectures we test, this transition produces qualitatively different failure modes: some models preserve substantial goal-conditioned behavior at vanishing attention, others fail despite carrying decodable goal information in their residual stream, and the depth at which this encoding emerges varies dramatically by architecture (from layer 2 to layer 27). A within-model causal ablation that closes the attention channel by force on Mistral collapses recall from near-perfect to eleven percent on a 20-fact retention task and raises persona-constraint violations to levels exceeding the adversarial-pressure baseline despite no user pressure, with both effects emerging at the predictable crossover turn. Linear probes on residual representations recover per-episode recall outcomes with AUC up to 0.99 across all four primary architectures (input embedding: chance), evidencing the second channel and showing its depth profile is architecture-specific. Across multiple model architectures and model scales, we show that the attention channel and the residual channel are separable, and that the gap between attention loss and residual capacity determines whether goal-conditioned behavior survives a long conversation. We provide GAR as a metric, the channel transition framework as a mechanism, and a parametric prediction of when multi-turn instruction-following will fail.

Learn More