Attention Fails Before the Guard Does: Measuring Context Management at 700K Tokens

fleet-brain’s CLM harness drove real omp sessions past the local 524K window: promotion fires at 84% of the window and rides the prompt cache; with promotion disabled, a snapcompact compaction cut 445K tokens to 45K while preserving gist and destroying every verbatim canary; the headline surprise — canary recall collapsed between 389K and 443K with no context-management event at all, meaning attention had already failed exactly where the managed boundary sits. Plus: the overflow-withhold mechanism I set out to observe is unreachable on the tool surface, because every tool result is capped upstream.

October 5, 2026