<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Self-Hosted on Bits-Entangled</title><link>https://stondo.github.io/tags/self-hosted/</link><description>Recent content in Self-Hosted on Bits-Entangled</description><generator>Hugo -- 0.157.0</generator><language>en-us</language><lastBuildDate>Tue, 25 Aug 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://stondo.github.io/tags/self-hosted/index.xml" rel="self" type="application/rss+xml"/><item><title>AIOS: a Three-Node Self-Hosted Agentic Coding Stack with DeepSeek V4 Flash, vLLM, and a Custom Model Router</title><link>https://stondo.github.io/posts/aios-three-node-agentic-coding-stack/</link><pubDate>Fri, 21 Aug 2026 00:00:00 +0000</pubDate><guid>https://stondo.github.io/posts/aios-three-node-agentic-coding-stack/</guid><description>How I wired three GPU machines (2x RTX PRO 6000, RTX 5090, RTX 4080) into a single agentic coding stack: SGLang serving DeepSeek V4 Flash as the brain, a Qwen3.8 27B NVFP4 on the 5090 as the workhorse for subagents and vision, a custom chain-based router for failover and code exploration, and persistent agent memory across opencode, pi, and qwen CLI.</description></item></channel></rss>