<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Posts on Bits-Entangled</title><link>https://stondo.github.io/posts/</link><description>Recent content in Posts on Bits-Entangled</description><generator>Hugo -- 0.162.1</generator><language>en-us</language><lastBuildDate>Wed, 26 Aug 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://stondo.github.io/posts/index.xml" rel="self" type="application/rss+xml"/><item><title>Zoysh: Yosh's yo Comes to zsh</title><link>https://stondo.github.io/posts/zoysh-yosh-zsh-port/</link><pubDate>Tue, 25 Aug 2026 13:00:00 +0000</pubDate><guid>https://stondo.github.io/posts/zoysh-yosh-zsh-port/</guid><description>Zoysh is a zsh port of Yosh, the LLM-enabled shell by Fil Pizlo. It generates shell commands from natural language, prefills them for review, streams answers, runs multi-step plans, and never executes anything you did not press Enter on. Works with local models by default.</description></item><item><title>Three Days, Four Wrong Hypotheses, and One Uncensored 27B That Finally Serves on a Single RTX 5090</title><link>https://stondo.github.io/posts/obliteratus-qwen38-27b-nvfp4-rtx5090/</link><pubDate>Tue, 25 Aug 2026 12:00:00 +0000</pubDate><guid>https://stondo.github.io/posts/obliteratus-qwen38-27b-nvfp4-rtx5090/</guid><description>The full story of quantizing OBLITERATUS/Qwen3.8-27B-OBLITERATED to NVFP4 for a single RTX 5090 with vLLM and LMCache: every ModelOpt trap I fell into, the assertion that survived three fixes, and the one debugging trick that cracked it in a single run.</description></item><item><title>AIOS: a Three-Node Self-Hosted Agentic Coding Stack with DeepSeek V4 Flash, vLLM, and a Custom Model Router</title><link>https://stondo.github.io/posts/aios-three-node-agentic-coding-stack/</link><pubDate>Fri, 21 Aug 2026 00:00:00 +0000</pubDate><guid>https://stondo.github.io/posts/aios-three-node-agentic-coding-stack/</guid><description>How I wired three GPU machines (2x RTX PRO 6000, RTX 5090, RTX 4080) into a single agentic coding stack: SGLang serving DeepSeek V4 Flash as the brain, a Qwen3.8 27B NVFP4 on the 5090 as the workhorse for subagents and vision, a custom chain-based router for failover and code exploration, and persistent agent memory across opencode, pi, and qwen CLI.</description></item><item><title>Why nvidia_peermem Fails on DGX Spark: a Deep Dive into GPUDirect RDMA, Secure Boot, and Unified Memory</title><link>https://stondo.github.io/posts/nvidia-peermem-gpudirect-rdma-dgx-spark/</link><pubDate>Mon, 23 Feb 2026 00:00:00 +0000</pubDate><guid>https://stondo.github.io/posts/nvidia-peermem-gpudirect-rdma-dgx-spark/</guid><description>A detailed investigation into why nvidia_peermem refuses to load on HP ZGX Nano G1n (DGX Spark) systems, covering Secure Boot module signing, kernel source code analysis, the fundamental hardware limitation that makes GPUDirect RDMA impossible on DGX Spark / GB10 SoC, and the PCIe Gen5 x4 bandwidth bottleneck on the ConnectX-7 NICs.</description></item></channel></rss>