Qwen3.8-Flash-Next at 512K Context on Two RTX PRO 6000: 146 t/s, a 51B-Entry Table in Host RAM, and the FP8 KV Door That Stays Shut

Flash-Next is Qwen’s sparse-attention flagship and its official FP8 recipe assumes datacenter GB300. Adapting it to 2x96 GB workstation Blackwell took one legit trick (offloading the 51B-entry PLE n-gram table to pinned host RAM, freeing ~25 GiB per GPU for KV), one config archaeology exercise (YaRN via –hf-overrides because the build has no –rope-scaling flag), and discipline about what not to retry: fp8 KV hard-requires BF16 on QSA and fails differently on vLLM and SGLang, and 786K+ contexts OOM during compile no matter the KV math. Final stack: 512K context, 1.42M-token pool, 146.5 tok/s decode with MTP ns=3, 9-11.2K tok/s prefill, vision and tools validated. The SGLang NVFP4 day-0 lane runs too (after three SM120 gates), but serves one request at 262K slower than the vLLM lane serves two at 512K, so vLLM keeps the port.

August 27, 2026

One Server, Every Mode: MiniMax H3 Goes From Static Frames to Full Video Editing on Two RTX PRO 6000

H3 ships as two checkpoints: FL2VA (text + first/last frame) and Ref2VA (references including video). Serving both DiTs from one process took three failed hypotheses, a basename-sensitive model resolver, a 66 GB minimal partition download, and one multipart field spelled plural. Now one systemd unit serves text-to-video, keyframe animation, and video editing at 56 GiB per GPU, validated end to end.

August 27, 2026

262,144 Tokens on One RTX 5090: Qwen3.8-27B with an NVFP4 KV Cache That Actually Works

My 27B coding brain was hard-capped at ~107K context by fp8 KV and fat weights. A stranger’s gist promised a 451K-token KV pool on the same GPU via NVFP4 KV cache and a patched vLLM. I ported it, pinned it, validated it with planted-password recall tests at 235K tokens — and then watched it crash spectacularly the first time a real agent used it. The final stack: full 262,144-token context, ~125 tok/s decode, vision, MTP speculative decoding, all on one card.

August 27, 2026

Zoysh: Yosh's yo Comes to zsh

A zsh plugin that ports Yosh’s yo interaction model: natural language in, reviewed command at your prompt, streaming answers, multi-step plans, scrollback context, and an experimental native module. Local-first, multi-provider, and safe by construction.

August 25, 2026

Three Days, Four Wrong Hypotheses, and One Uncensored 27B That Finally Serves on a Single RTX 5090

No NVFP4 build of the uncensored Qwen3.8-27B existed, so I made one. It took a fake-quantization scare, a rename odyssey across 800 tensors, an MTP file that quietly loaded the vision tower twice, and one instrumented container run that proved everything I believed was wrong. 40 tok/s single stream, 317 tok/s with 8 concurrent subagents.

August 25, 2026

AIOS: a Three-Node Self-Hosted Agentic Coding Stack with DeepSeek V4 Flash, vLLM, and a Custom Model Router

Instead of picking one model per task, I made three machines cooperate: a big thinking model for primary reasoning, a fast local 27B for subagent loops and images, a router with automatic failover in front of everything, and shared agent memory. Total cost: electricity. Here is the full architecture, the benchmarks, and the lessons learned the hard way.

August 21, 2026

Why nvidia_peermem Fails on DGX Spark: a Deep Dive into GPUDirect RDMA, Secure Boot, and Unified Memory

I chased a modprobe error across two DGX Spark nodes, peeled back layers of Secure Boot signing and kernel build flags, read the nvidia_peermem source code, and discovered that GPUDirect RDMA is architecturally unsupported on the GB10 Grace Blackwell SoC. Here’s the full story.

February 23, 2026 · stondo