<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Sparse-Attention on Bits-Entangled</title><link>https://stondo.github.io/tags/sparse-attention/</link><description>Recent content in Sparse-Attention on Bits-Entangled</description><generator>Hugo -- 0.162.1</generator><language>en-us</language><lastBuildDate>Thu, 27 Aug 2026 22:00:00 +0000</lastBuildDate><atom:link href="https://stondo.github.io/tags/sparse-attention/index.xml" rel="self" type="application/rss+xml"/><item><title>Qwen3.8-Flash-Next at 512K Context on Two RTX PRO 6000: 146 t/s, a 51B-Entry Table in Host RAM, and the FP8 KV Door That Stays Shut</title><link>https://stondo.github.io/posts/qwen38-flash-next-512k-two-rtx-pro-6000/</link><pubDate>Thu, 27 Aug 2026 22:00:00 +0000</pubDate><guid>https://stondo.github.io/posts/qwen38-flash-next-512k-two-rtx-pro-6000/</guid><description>The official vLLM recipe for Qwen3.8-Flash-Next trusts GB300 at TP4. I wanted it on two workstation-class RTX PRO 6000 cards: TP2, the PLE n-gram table offloaded to pinned host RAM, YaRN stretched to 512K context, MTP speculative decoding. That stack delivers a 1.42M-token KV pool, 146.5 tok/s decode and working vision. It also taught me that fp8 KV is architecturally dead on QSA (two engines, two failure modes), that systemd strips quotes out of JSON args, and that 1M context dies at compile, not at KV budgeting.</description></item></channel></rss>