Terminal-style card: local-first AI model routing across three boxes and two cloud providers

One Coding Agent, Three GPUs, Two Cloud Keys: Local-First Routing and the Judge That Was Quietly Burning My 5090

An audit of my coding agent’s model routing turned into a routing philosophy plus three real bugs. The principle: route by token volume, not vibes — the default loop, subagent fan-out, and background work run on local GPUs for free, while planning, review, and deep reasoning buy cloud quality. Found along the way: a context-promotion target pointing at a provider I have no credential for (a silent no-op), omp’s internal judge role degrading to prompted 27B judging because no TypeSafe key exists (fixed with Ollaya’s winnow:e4b decision model on the senses node, calibrated answers in 0.93 s on CPU), and the DeepSeek V4.1 naming trap where the literal model id does not exist — deepseek-flash IS V4.1-Flash, and deepseek-v4-pro server-routes to it anyway. Plus the GLM-5.3 flash/base vision trap, verified with live image requests.

October 2, 2026

Zoysh: Yosh's yo Comes to zsh

A zsh plugin that ports Yosh’s yo interaction model: natural language in, reviewed command at your prompt, streaming answers, multi-step plans, scrollback context, and an experimental native module. Local-first, multi-provider, and safe by construction.

August 25, 2026

AIOS: a Three-Node Self-Hosted Agentic Coding Stack with DeepSeek V4 Flash, vLLM, and a Custom Model Router

Instead of picking one model per task, I made three machines cooperate: a big thinking model for primary reasoning, a fast local 27B for subagent loops and images, a router with automatic failover in front of everything, and shared agent memory. Total cost: electricity. Here is the full architecture, the benchmarks, and the lessons learned the hard way.

August 21, 2026