262,144 Tokens on One RTX 5090: Qwen3.8-27B with an NVFP4 KV Cache That Actually Works

My 27B coding brain was hard-capped at ~107K context by fp8 KV and fat weights. A stranger’s gist promised a 451K-token KV pool on the same GPU via NVFP4 KV cache and a patched vLLM. I ported it, pinned it, validated it with planted-password recall tests at 235K tokens — and then watched it crash spectacularly the first time a real agent used it. The final stack: full 262,144-token context, ~125 tok/s decode, vision, MTP speculative decoding, all on one card.

August 27, 2026

Three Days, Four Wrong Hypotheses, and One Uncensored 27B That Finally Serves on a Single RTX 5090

No NVFP4 build of the uncensored Qwen3.8-27B existed, so I made one. It took a fake-quantization scare, a rename odyssey across 800 tensors, an MTP file that quietly loaded the vision tower twice, and one instrumented container run that proved everything I believed was wrong. 40 tok/s single stream, 317 tok/s with 8 concurrent subagents.

August 25, 2026

AIOS: a Three-Node Self-Hosted Agentic Coding Stack with DeepSeek V4 Flash, vLLM, and a Custom Model Router

Instead of picking one model per task, I made three machines cooperate: a big thinking model for primary reasoning, a fast local 27B for subagent loops and images, a router with automatic failover in front of everything, and shared agent memory. Total cost: electricity. Here is the full architecture, the benchmarks, and the lessons learned the hard way.

August 21, 2026