My coding agent spends most of its life talking to hardware I own. That was the design intent, anyway. So when I sat down to answer one question — is my agent config actually using my local AI in a meaningful way, or is it decoration? — I expected a config review. What I got was a routing philosophy check plus three genuine bugs, one of which had been quietly taxing my GPU on every single turn.
The agent is omp (oh-my-pi, can1357’s fork of Pi), and the hardware is the fleet I described in the three-node agentic stack post: a two-GPU brain, a workhorse, and a senses node. Nothing here is exotic. What follows is the routing that came out of the audit, and the traps I hit while verifying it.
The fleet, in one paragraph
The brain is a pair of RTX PRO 6000s serving qwen3.8-flash-next at 512K
context through vLLM — the box I built for half-a-million-context
prefills. The workhorse
is an RTX 5090 running my NVFP4-KV quantization of Qwen3.8-27B — the one from
the 262K single-card build
and the OBLITERATUS
quantization saga — which
handles eight concurrent subagents without breaking a sweat. The senses
node is an RTX 4080 running SearXNG, embeddings, a reranker, speech in and
out, and the monitoring stack. On top of that, exactly two cloud API keys:
z.ai for GLM 5.3, and DeepSeek.
Local compute for volume. Cloud for decisions. That’s the whole thesis — the interesting part is making an agent harness actually respect it.
Route by token volume, not vibes
The mistake I see in most “hybrid” setups is routing by task type (“coding is local, research is cloud”) instead of by token economics. The right question for each role is: where does the volume go, and where does the stakes-per-token go?
omp has nine model roles plus a few model-kind roles, which makes this concrete. Here’s the split after the audit:
| Role | What it does | Token volume | Runs on |
|---|---|---|---|
default | every turn, every file read, all tool results | ~70-80% of session tokens | brain, local |
task / smol | subagent fan-out, summaries | bursty, huge | 5090, local |
tiny / memory / commit | titles, memory extraction, commit messages | high-frequency, small | 5090, local |
judge | internal agent judgments, every turn | tiny but constant | senses node, local |
web | all searches | frequent | senses node, local |
plan / advisor | planning, review, oversight | rare, critical | GLM 5.3, cloud |
slow | deep reasoning | rare, heavy | DeepSeek, cloud |
vision | screenshots, diagrams | rare | glm-5.3-flash, cloud |
The single most expensive thing an agent does — a 400K-token prefill on the default loop — happens on hardware I already paid for. Cloud spend scales with decisions, not with context. That asymmetry is the entire game.
Two timeouts make the local-first part practical, and both were learned the hard way:
providers:
streamFirstEventTimeoutSeconds: 600 # local prefill of 400k+ prompts
streamIdleTimeoutSeconds: 120 # long reasoning streams between tokens
If you’ve ever watched a local vLLM chew through a giant prefill (I wrote about a related prefill footgun in the zoysh port), you know the default first-event timeouts in agent harnesses assume cloud latencies. They don’t survive contact with a 512K-context local box.
Bug 1: the dead promotion target
omp has a lovely feature called context promotion: when a turn overflows the context window, instead of compacting, it can switch to a designated larger-context model and retry. My deep model had:
contextPromotionTarget: "anthropic/claude-opus-5.5"
Sounds great. One problem: there is no Anthropic credential on this machine. The promotion target is availability-checked at overflow time, and a target that can’t resolve a credential is silently skipped — every overflow went straight to compaction instead. A one-line config value, dead on arrival, failing quietly forever.
The fix was to point it at something I actually have a key for, with a bigger window:
contextPromotionTarget: "deepseek/deepseek-v4-pro" # 1M ctx, key present
Lesson: any “fall back to X on failure” config is a liability until you’ve watched it fail at least once. Dead references don’t error — they no-op.
Bug 2: the judge that was quietly burning my 5090
This is my favorite find of the audit. omp makes small typed decisions about
its own state — classifying a prompt to pick a thinking-effort level,
detecting when a stream stopped for unexpected reasons, AI-staging files in
the git UI. These run through a judge role with a native chain that starts
at TypeSafe’s System One models and degrades through progressively cheaper
options.
I have no TypeSafe key. So every one of those judgments — on every turn — was
falling through to the tiny role: a prompted 27B LLM generating tokens to
answer yes/no questions, uncalibrated, on the same GPU that runs my
subagent fan-out. The config looked innocent. The GPU bill was real.
The fix is a project I’d been waiting for an excuse to deploy:
Ollaya, a local server for decision models —
single-forward-pass models that return calibrated probabilities instead of
generated prose. And the reason it drops straight into omp is a small
miracle of protocol archaeology: omp’s models.yml supports api: typesafe
as a custom provider API — the exact wire format (/v1/systemone) Ollaya
speaks.
# models.yml
ollaya:
baseUrl: http://senses-node:11435
api: typesafe # System One judgment wire -> the judge role
auth: none
models:
- id: winnow:e4b
name: Winnow E4B (local decision model)
contextWindow: 8192
maxTokens: 1024
# config.yml
modelRoles:
judge: ollaya/winnow:e4b
Live test, from the same request shape omp sends:
{"root_cause": {"choice": "db_outage", "confidence": 0.957,
"probabilities": {"db_outage": 0.97, "app_bug": 0.027, "config": 0.002}},
"is_urgent": {"choice": "yes", "confidence": 0.973}}
0.93 seconds warm — on CPU, because the 4080’s VRAM belongs to the embeddings and reranker services (on GPU it’s ~90 ms). Even on CPU it beats prompted-27B judging on latency, returns actual calibrated probabilities, and costs the 5090 exactly nothing. The judge role now runs on a box whose job is already “the senses.”
The V4.1 naming trap
“Does the DeepSeek API offer V4.1?” Yes. No — yes, but hear me out.
Every catalog names it differently. omp’s bundled catalog showed
deepseek-v4-pro and deepseek-flash with no v4.1 anywhere. OpenRouter had
a literal deepseek-v4.1-flash. HuggingFace had DeepSeek-V4.1-Flash. So I
stopped trusting catalogs and asked the endpoint:
$ curl https://api.deepseek.com/v1/models -H "Authorization: Bearer $KEY"
{"data":[
{"id":"deepseek-flash","name":"DeepSeek-V4.1-Flash", ...},
{"id":"deepseek-v4-pro","name":"DeepSeek-V4-Pro", ...}]}
Ground truth: there is no literal v4.1 model id on the DeepSeek API.
deepseek-flash is V4.1-Flash — 1M context, vision-capable, rolling
alias. And the kicker from their release notes: since September 14, all
deepseek-v4-pro requests are server-side routed to V4.1-Flash at Flash
prices, “until V4.1-Pro launches.”
Which means my slow role — set to deepseek-v4-pro:high — is already
serving V4.1-Flash, and when V4.1-Pro ships it will auto-upgrade to the real
flagship without me touching anything. Sometimes the lazy selector is the
correct one.
The transferable lesson is older than LLMs: GET /v1/models on the live
endpoint beats every catalog. Catalogs drift; endpoints don’t lie about
what they serve.
Vision: the flash/base trap
One more verified-fact detour, because it’s a trap wearing a friendly name. GLM-5.3 — the big flagship — is text-only. GLM-5.3-Flash is the multimodal one: native vision encoder, 1M context, image/video/file input. If you skim model names the way I do, “use the bigger model for vision” is exactly the mistake you’d make.
And the local option? At audit time, no. My 27B workhorse is not a VLM, and the receipt still stands — a live image request returning:
HTTP 400: "At most 0 image(s) may be provided in one prompt."
So vision lives on zai/glm-5.3-flash:high, which is genuinely good at it
(screenshot-to-app understanding is apparently a design goal) and comes with
3× quota on the coding plan. Test your assumptions with actual requests; the
config comment that says “not a VLM” might be stale, and the one that says
nothing might be wrong too.
This section then wrote its own epilogue: when I automated exactly that test
(see the P.S.), the red-pixel probe came back from the deep box’s
qwen3.8-flash-next with “a soft red or pink” — a correct answer from a
vLLM build that had quietly been multimodal all along. “The local option
doesn’t exist” was the same categorical mistake this section warns about,
made by the author of the section. Model names tell you nothing about the
build behind them — query every endpoint, including the ones you are sure
about.
And the routing followed the receipt: vision still leads with the cloud
model — quality first, and vision volume is low — but its fallback chain is
now all-VLM, with the local flash-next as hop one. Five identical
consecutive probes name the color in 0.3 s on hardware I own.
The lattice
Fallback chains are where local-first gets to be clever, because the chain can cross the local/cloud boundary in both directions:
retry:
fallbackChains:
deep/*: [5090-27b, zai/glm-5.3] # local -> local -> cloud
fast/*: [deep/flash-next, zai/glm-5.3] # local -> local -> cloud
zai/*: [deepseek/deepseek-v4-pro, deep] # cloud -> other cloud -> local
deepseek/*: [deep/flash-next, zai/glm-5.3] # cloud -> local -> other cloud
A cloud outage degrades to my hardware. A hardware reboot degrades to the other cloud. No single failure takes the agent down.
One honest footnote — since fixed. The zai/* chain technically serves
vision too, and when this post was published its fallbacks were text-only
models: a z.ai 429 mid-screenshot-analysis would have landed somewhere that
couldn’t see the image. The day the deep box proved itself a real VLM (next
section), vision got its own role-scoped chain — local flash-next first,
deepseek-flash second — so every hop of it can actually see. A z.ai outage
now degrades to hardware I own instead of a model that’s blind.
Two kinds of memory
The Context Language Models
extension went in during the audit and turned out to be the interesting
compatibility surprise: it’s built against upstream Pi’s @earendil-works/*
packages, while omp is the @oh-my-pi/* fork — and omp’s plugin installer
resolves those peer dependencies anyway. Verified live: the extension’s
live_context_annotate and live_context_recall tools register in every
fresh session, and a /clm panel shows the budget the overflow guard is
enforcing. (License note, corrected from my first draft: the extension itself
is MIT; only Facebook’s reference implementation carries CC BY-NC.)
Which raises the obvious question: I already run Cairnkeep, my durable memory layer, wired into the same agent. Do they fight?
No — they’re different layers with different lifetimes:
- CLM is working memory. It curates what the model can see right now, inside this session’s window: withhold the oldest tool results near budget, let the model edit its own context file, keep prefills lean.
- Cairnkeep is durable memory. It holds what must survive after the session — accepted decisions, pitfalls, root causes — project-scoped, cross-harness, retrieval-first, and never an authority over the repo.
But there is one real seam, and it’s the kind that bites quietly: the more aggressively you prune the window, the more it matters that everything durable was already persisted. A fat context hides sloppy memory hygiene — the model “remembers” because the transcript is still in view. A CLM-curated context doesn’t forgive that. When the guard withholds an observation, it leaves a note pointing at the saved file — and re-reading that file is a fresh prefill.
The integration, then, is not code. Both surfaces are already model-facing
tools, so the whole thing is one steering document — the extension’s
PI_CLM_STEERING appends a policy to the system prompt, and mine encodes
four rules:
1. Persist before pruning — a durable conclusion about to leave the window
gets memory_write'd first.
2. Recall before re-reading — memory_search beats re-prefilling a saved file.
3. Memory is context, never authority — a memory that conflicts with the
repo, a test, or a config file loses.
4. Distill — no raw transcripts into durable memory.
Working memory forgets on purpose. Durable memory is the part that’s supposed to stick. Point each at its own job and they compose instead of compete.
What’s next
The audit left a queue: point omp’s memory backend at mnemopi with the local
embeddings server doing the vectors; offload the tiny roles to omp’s
embedded tiny models and raise subagent concurrency; and wire the senses
node’s jury and dispatch MCPs into the agent.
Verdict
Is the config using local AI in a meaningful way? After the audit, honestly: yes. Local carries the daily driving, all fan-out, all background machinery, judging, and search; cloud is reserved for exactly the four places quality justifies it. The bugs weren’t in the philosophy — they were in the verification. Dead promotion targets, silent judge degradation, and phantom model names all share one cure:
Don’t read the config. Query the endpoints.
P.S. (2026-10-04): “Query the endpoints” is a command now
The line above shipped as a tool two days later: fleet-doctor, a small
Python CLI that reads the same config.yml/models.yml omp reads and does
four things. probe checks endpoint health, catalog-vs-roster drift,
credential resolution for every promotion and fallback target, live vision
receipts, and judge reachability. bench produces TTFT-vs-context curves per
tier plus a cost-per-role table (local tiers read free; the cloud meters
tick in the same table). chaos runs induced-failure drills — dry-run is the
default, a live run needs two flags, every drill ships with a documented
recovery. judge asks the decision model ad-hoc questions directly. Exit
codes are 0 clean / 1 degraded / 2 broken, and a systemd timer runs the
check every 30 minutes and feeds the series to the existing Grafana.
Two receipts worth the price of admission. The audit bugs are now regression
tests: a planted ghost promotion target exits 2, a planted -old model id
behind a catalog claim exits 1, and the suite stays red until the tools learn
to catch them.
And the better story: the first live stop-service drill passed —
systemctl stop returned 0, the verification probe stayed green, recovery
looked trivial. It was a lie. The unit I stopped was a same-named
system-scope decoy; the live vLLM container belonged to a user-scope
quadlet, and its orphaned conmon kept serving through the whole “drill”. The
tool caught it precisely because it verifies the endpoint and not the
command’s exit status. Re-run with the right lever: real teardown, connection
reset, measured cold reload — MTTR 236 s — while the coding agent riding
that very endpoint failed over from the deep box to the 5090 mid-drill and
kept working.
The lesson the tool exists to enforce, re-learned by the tool: a green check is a claim, not a fact. Query the endpoints. Then query them again after you pull the lever.
Final coda, same day: the tool’s own vision probe spent its first timer runs
reporting the deep box as unsure — accepted image, no color in the answer.
Not model flakiness; max_tokens: 16, swallowed whole by the reasoning
prefix (finish=length, empty content) before the answer was ever emitted.
Raise the cap to 512, make the probe image non-degenerate, pin both in a
regression test — and the receipt reads verified, five for five, at 0.3 s.
Even the verifier needed the lecture: a green check is a claim, not a fact.
Including mine.
