<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Ai-Agents on Bits-Entangled</title><link>https://stondo.github.io/tags/ai-agents/</link><description>Recent content in Ai-Agents on Bits-Entangled</description><generator>Hugo -- 0.162.1</generator><language>en-us</language><lastBuildDate>Mon, 05 Oct 2026 15:30:00 +0200</lastBuildDate><atom:link href="https://stondo.github.io/tags/ai-agents/index.xml" rel="self" type="application/rss+xml"/><item><title>Attention Fails Before the Guard Does: Measuring Context Management at 700K Tokens</title><link>https://stondo.github.io/posts/clm-context-pressure-measured/</link><pubDate>Mon, 05 Oct 2026 15:30:00 +0200</pubDate><guid>https://stondo.github.io/posts/clm-context-pressure-measured/</guid><description>Everyone has opinions about long-context agents; nobody had measurements of what the machinery actually does under sustained pressure on a local fleet. I built a harness that drives a real coding-agent session past a 524K-token model window, unattended, and measured five things: where promotion fires, what compaction destroys, when verbatim recall collapses, why the truncation guard I was testing never fired, and how the whole story reorders my assumptions about agent memory hygiene.</description></item><item><title>One Coding Agent, Three GPUs, Two Cloud Keys: Local-First Routing and the Judge That Was Quietly Burning My 5090</title><link>https://stondo.github.io/posts/omp-local-first-routing-homelab-fleet/</link><pubDate>Fri, 02 Oct 2026 09:00:00 +0000</pubDate><guid>https://stondo.github.io/posts/omp-local-first-routing-homelab-fleet/</guid><description>I audited the model routing of my coding agent over the three-box fleet and found one dead fallback, one silent GPU tax, and one model name that does not exist. The fix routing: local models carry the token volume, cloud keys carry the decisions, and a 0.9-second local decision model now does every internal judgment.</description></item><item><title>Top 12% at the AI Chessathon: One CPU Core, Four Dead Neural Nets, and the Residual That Finally Earned Its Elo</title><link>https://stondo.github.io/posts/aichessathon-numba-nnue-engine-58-of-465/</link><pubDate>Thu, 24 Sep 2026 09:00:00 +0000</pubDate><guid>https://stondo.github.io/posts/aichessathon-numba-nnue-engine-58-of-465/</guid><description>I entered an online tournament where AIs play chess: 465 bots, one CPU core, no GPU, no network, a 50 MB zip, and 120 seconds per game. The engine that finished 58th is a numba-jitted alpha-beta in pure Python with a 128-wide NNUE residual on top of a classical evaluation. Four neural networks died to teach me that lower training loss does not play chess.</description></item><item><title>Announcing Cairnkeep: Durable, Local-First Memory for Coding Agents, Without the Autonomy</title><link>https://stondo.github.io/posts/announcing-cairnkeep-durable-memory-coding-agents/</link><pubDate>Thu, 27 Aug 2026 23:00:00 +0000</pubDate><guid>https://stondo.github.io/posts/announcing-cairnkeep-durable-memory-coding-agents/</guid><description>Every coding agent I run forgets everything when the session ends, every harness silos its own memory, and most agent-memory projects gradually turn into autonomous agents that write and approve things on their own. Cairnkeep is my answer: a durable, local-first memory and context layer for coding agents that works across Claude Code, OpenCode, Codex, Kimi, Qwen and Pi, ships a 26-lesson release-verified curriculum, and is built around one stubborn rule: memory is durable context, never an authority.</description></item></channel></rss>