Attention Fails Before the Guard Does: Measuring Context Management at 700K Tokens
fleet-brain’s CLM harness drove real omp sessions past the local 524K window: promotion fires at 84% of the window and rides the prompt cache; with promotion disabled, a snapcompact compaction cut 445K tokens to 45K while preserving gist and destroying every verbatim canary; the headline surprise — canary recall collapsed between 389K and 443K with no context-management event at all, meaning attention had already failed exactly where the managed boundary sits. Plus: the overflow-withhold mechanism I set out to observe is unreachable on the tool surface, because every tool result is capped upstream.

One Coding Agent, Three GPUs, Two Cloud Keys: Local-First Routing and the Judge That Was Quietly Burning My 5090
An audit of my coding agent’s model routing turned into a routing philosophy plus three real bugs. The principle: route by token volume, not vibes — the default loop, subagent fan-out, and background work run on local GPUs for free, while planning, review, and deep reasoning buy cloud quality. Found along the way: a context-promotion target pointing at a provider I have no credential for (a silent no-op), omp’s internal judge role degrading to prompted 27B judging because no TypeSafe key exists (fixed with Ollaya’s winnow:e4b decision model on the senses node, calibrated answers in 0.93 s on CPU), and the DeepSeek V4.1 naming trap where the literal model id does not exist — deepseek-flash IS V4.1-Flash, and deepseek-v4-pro server-routes to it anyway. Plus the GLM-5.3 flash/base vision trap, verified with live image requests.
Top 12% at the AI Chessathon: One CPU Core, Four Dead Neural Nets, and the Residual That Finally Earned Its Elo
A tournament report from the AI Chessathon: a pure-Python numba chess engine on one CPU core, four replacement neural evaluators that all lost to classical piece-square tables (one 0-30), the residual NNUE that beat them (+224 Elo, confirmed at tournament time control), a time-management bug found by post-mortem forensics mid-event, and the promotion gates that kept an AI coding agent honest. Final: 58th of 465, rating 2207.
Every Flag Justified: The Complete DeepSeek-V4.1-Flash Lane Recipe for Two RTX PRO 6000
A copy-paste serving recipe for DeepSeek-V4.1-Flash at EXL3 2bpw on 2× RTX PRO 6000 Blackwell (sm_120): rootless podman over the diffbot pack, eager mode forced by the Engram-on-NVMe path, the generation_config.json that ships official sampling to every client, the reasoning-effort knob that wants strings and not integers, and the measured performance envelope of the result.
The Word Salad Only opencode Could See: Hunting a Phantom Garbage Bug on DeepSeek-V4.1-Flash
CLI-dependent degradation on a quantized 552B MoE: concurrency hammers, a logging proxy that exposed what opencode really sends (nothing — and that was the bug), one reproduced episode of fluent Chinese nonsense in twelve cycles, and the missing generation_config.json that separated an uncapped top_p 1.0 tail from DeepSeek’s official recipe. Plus the reasoning-effort knob that exists as 0-100 in one engine and string levels in another.
552B Parameters and the 203 GB You Keep on Disk: DeepSeek-V4.1-Flash on Two RTX PRO 6000
The quant landscape sweep that ended at a 2.0bpw EXL3 pack built for exactly two RTX PRO 6000s, why 196B of Engram tables live on NVMe in every configuration that exists, what eager-mode serving actually measures (94-104 tok/s decode, 2.7K tok/s prefill), and the AM5 ceiling that keeps the last 15% of the recipe’s performance behind a RAM upgrade the platform cannot physically provide.
The 30-Second Watchdog That Ate a Four-Hour Render: Long-Form MiniMax H3 on Two RTX PRO 6000
A 15s arcade-to-live-action video edit on H3 failed four different ways before it succeeded: cudnn graph crashes on long sequences, a 30s async-output watchdog that throws away finished renders, flashinfer kernels that require tcgen05 MMA my RTX PRO 6000s don’t have, and a finisher script that reported failure after succeeding. The winning recipe: patch the watchdog via bind-mount, CUDNN_ATTN, 540p reference, 1h45m end to end.
An Uncensored 305B With Eyes: DeepSeek-V4-Flash-Vision on 2× RTX PRO 6000, and the Two-Engine Week That Tamed It
The uncensored build of DeepSeek’s Vision-Exp is a byte-for-byte drop-in — which meant every failure afterwards was mine to fix. A sparse-MLA kernel that rejects vision prefill on SM120, a streaming tool-call parser that corrupted exactly the arguments coding agents live on, and a health check that proved the port but not the model. 142 tok/s single-stream, native vision, zero refusals, one new router feature.
From locklocklock to 262K Context: GLM-5.3-Flash on Two RTX PRO 6000
A 320B sparse-hybrid MoE serving word salad at temperature 0, a repro whose bugs outlived the bug it hunted, a per-layer bisect that cleared the KDA layers and convicted the sparse full-attention path, and an EXL3/TR3 plus B12X deployment that ends the story with verified 262K-context retrieval at 39 tok/s.
Announcing Cairnkeep: Durable, Local-First Memory for Coding Agents, Without the Autonomy
I have been building Cairnkeep since early July: project-scoped memory that survives sessions and crosses harnesses, a retrieval-first protocol agents must follow instead of new autonomy, least-authority MCP surfaces (read-only tool profiles, capability gates, immutable context packs), and optional, explicit opt-ins for everything networked. This post is the what and the why, the production deployment I run it in (a VPS-hosted memory server for a four-machine fleet, with the work laptop deliberately isolated), the 26-lesson learning path with its course labs repo, and the companion video series on my BitsEntangled YouTube channel.