<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Llm on Bits-Entangled</title><link>https://stondo.github.io/tags/llm/</link><description>Recent content in Llm on Bits-Entangled</description><generator>Hugo -- 0.162.1</generator><language>en</language><lastBuildDate>Tue, 25 Aug 2026 13:00:00 +0000</lastBuildDate><atom:link href="https://stondo.github.io/tags/llm/index.xml" rel="self" type="application/rss+xml"/><item><title>Zoysh: Yosh's yo Comes to zsh</title><link>https://stondo.github.io/posts/zoysh-yosh-zsh-port/</link><pubDate>Tue, 25 Aug 2026 13:00:00 +0000</pubDate><guid>https://stondo.github.io/posts/zoysh-yosh-zsh-port/</guid><description>Zoysh is a zsh port of Yosh, the LLM-enabled shell by Fil Pizlo. It generates shell commands from natural language, prefills them for review, streams answers, runs multi-step plans, and never executes anything you did not press Enter on. Works with local models by default.</description></item><item><title>AIOS: a Three-Node Self-Hosted Agentic Coding Stack with DeepSeek V4 Flash, vLLM, and a Custom Model Router</title><link>https://stondo.github.io/posts/aios-three-node-agentic-coding-stack/</link><pubDate>Fri, 21 Aug 2026 00:00:00 +0000</pubDate><guid>https://stondo.github.io/posts/aios-three-node-agentic-coding-stack/</guid><description>How I wired three GPU machines (2x RTX PRO 6000, RTX 5090, RTX 4080) into a single agentic coding stack: SGLang serving DeepSeek V4 Flash as the brain, a Qwen3.8 27B NVFP4 on the 5090 as the workhorse for subagents and vision, a custom chain-based router for failover and code exploration, and persistent agent memory across opencode, pi, and qwen CLI.</description></item></channel></rss>