<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Glm-5.3-Flash on Bits-Entangled</title><link>https://stondo.github.io/tags/glm-5.3-flash/</link><description>Recent content in Glm-5.3-Flash on Bits-Entangled</description><generator>Hugo -- 0.162.1</generator><language>en-us</language><lastBuildDate>Sun, 30 Aug 2026 09:00:00 +0000</lastBuildDate><atom:link href="https://stondo.github.io/tags/glm-5.3-flash/index.xml" rel="self" type="application/rss+xml"/><item><title>From locklocklock to 262K Context: GLM-5.3-Flash on Two RTX PRO 6000</title><link>https://stondo.github.io/posts/glm-5.3-flash-two-rtx-pro-6000-from-garbage-to-verified/</link><pubDate>Sun, 30 Aug 2026 09:00:00 +0000</pubDate><guid>https://stondo.github.io/posts/glm-5.3-flash-two-rtx-pro-6000-from-garbage-to-verified/</guid><description>GLM-5.3-Flash booted on my two RTX PRO 6000 Blackwell and answered every prompt with deterministic garbage. The investigation ran from a fake kernel bug I wrote myself, through a full exoneration of the linear-attention stack, to a live per-layer bisect that cornered the real culprit in the sparse-MLA prefill path. The fix came from a completely different quantization and kernel stack, and the model now passes 261,900-token needle retrieval on my desk.</description></item></channel></rss>