<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Bits-Entangled</title><link>https://stondo.github.io/</link><description>Recent content on Bits-Entangled</description><generator>Hugo -- 0.157.0</generator><language>en-us</language><lastBuildDate>Tue, 25 Aug 2026 13:00:00 +0000</lastBuildDate><atom:link href="https://stondo.github.io/index.xml" rel="self" type="application/rss+xml"/><item><title>Porting Yosh's yo to zsh in One Long Day: How Zoysh Was Born</title><link>https://stondo.github.io/posts/zoysh-yosh-zsh-port/</link><pubDate>Tue, 25 Aug 2026 13:00:00 +0000</pubDate><guid>https://stondo.github.io/posts/zoysh-yosh-zsh-port/</guid><description>How I fell in love with Yosh, the LLM-enabled shell built by Fil Pizlo, and ported its yo command to zsh as a plain script plugin: streaming answers, Ctrl-C cancellation, multi-step plans, a ZLE widget, scrollback capture, and an experimental native module. Plus the print -z footgun I shipped first and fixed honestly.</description></item><item><title>Three Days, Four Wrong Hypotheses, and One Uncensored 27B That Finally Serves on a Single RTX 5090</title><link>https://stondo.github.io/posts/obliteratus-qwen38-27b-nvfp4-rtx5090/</link><pubDate>Tue, 25 Aug 2026 12:00:00 +0000</pubDate><guid>https://stondo.github.io/posts/obliteratus-qwen38-27b-nvfp4-rtx5090/</guid><description>The full story of quantizing OBLITERATUS/Qwen3.8-27B-OBLITERATED to NVFP4 for a single RTX 5090 with vLLM and LMCache: every ModelOpt trap I fell into, the assertion that survived three fixes, and the one debugging trick that cracked it in a single run.</description></item><item><title>AIOS: a Three-Node Self-Hosted Agentic Coding Stack with DeepSeek V4 Flash, vLLM, and a Custom Model Router</title><link>https://stondo.github.io/posts/aios-three-node-agentic-coding-stack/</link><pubDate>Fri, 21 Aug 2026 00:00:00 +0000</pubDate><guid>https://stondo.github.io/posts/aios-three-node-agentic-coding-stack/</guid><description>How I wired three GPU machines (2x RTX PRO 6000, RTX 5090, RTX 4080) into a single agentic coding stack: SGLang serving DeepSeek V4 Flash as the brain, a Qwen3.8 27B NVFP4 on the 5090 as the workhorse for subagents and vision, a custom chain-based router for failover and code exploration, and persistent agent memory across opencode, pi, and qwen CLI.</description></item><item><title>Why nvidia_peermem Fails on DGX Spark: a Deep Dive into GPUDirect RDMA, Secure Boot, and Unified Memory</title><link>https://stondo.github.io/posts/nvidia-peermem-gpudirect-rdma-dgx-spark/</link><pubDate>Mon, 23 Feb 2026 00:00:00 +0000</pubDate><guid>https://stondo.github.io/posts/nvidia-peermem-gpudirect-rdma-dgx-spark/</guid><description>A detailed investigation into why nvidia_peermem refuses to load on HP ZGX Nano G1n (DGX Spark) systems, covering Secure Boot module signing, kernel source code analysis, the fundamental hardware limitation that makes GPUDirect RDMA impossible on DGX Spark / GB10 SoC, and the PCIe Gen5 x4 bandwidth bottleneck on the ConnectX-7 NICs.</description></item></channel></rss>