In the previous H3 post I got every MiniMax H3 mode serving from one systemd unit, with 4-second smoke tests as the proof. Smokes are lovely because they fit in a coffee break. The first real job I queued — a 15-second Ref2VA edit, my 1986 OutRun arcade clip morphing into photorealistic live action halfway through — sampled for 4 hours 21 minutes, finished, and was then thrown away at the finish line by a hardcoded 30-second timeout that appears in no documentation.
This is the autopsy, the kernel hunt that followed, and the rerun that made it: same job, 1h45m end to end, video embedded below.
The job
One prompt, one reference clip (357 frames at 24 fps, 14.875s), task ref2va: keep the first 7 seconds in the original sprite-based pixel-art style, then between seconds 7 and 8.5 melt the pixels into a real red 1980s convertible on a real coastal highway, and stay photorealistic to the end. Output is the shape H3 always emits: 1344×768 at 24 fps, 362 frames, H.264.
Yesterday’s run did all of that correctly — the denoise loop reached step 49/49, the VAE decoded, and then the orchestrator reported success into a void. inference_time_s: 15696, empty error message, no file.
Autopsy: three blockers, one fatal
Blocker 1: cudnn attention crashes on long sequences. The quadlet defaulted to CUDNN_ATTN, validated only on 4s smokes. At 362 frames the cudnn SDPA graph fails with mha_graph.execute ... got false, then a poisoned CUDA context throws illegal-memory-access on everything afterwards. I had flipped the unit to TORCH_SDPA as the safe dispatcher fallback. It works. It is also ~6-7x slower, which is why the sampling took 4h21m.
Blocker 2: reference duration is validated strictly. A 15.083s reference is rejected (duration must be in [2, 15] seconds). The source was 60 fps; the fix was a resample to 24 fps and a trim to 14.875s. Minor, but it cost a submission round-trip.
Blocker 3, the killer: a 30-second watchdog on the output path. Buried in vllm_omni/diffusion/diffusion_engine.py:
_ASYNC_OUTPUT_TIMEOUT = 30.0 # seconds
After the denoise completes, the post-processing (VAE tile decode of 362 frames plus device-to-host copy) runs as an async output. On a fast pod this fits in 30 seconds and the constant never bites. On my rig, on the slow SDPA path, it does not — asyncio.wait_for fires, the orchestrator aborts, and since the MP4 is encoded exactly once at the very end with no partial flush, the entire render is discarded. Not env-configurable. Not documented. Just a module-level constant with an opinion about how fast your GPU should be.
The kernel hunt on sm_120a
Before rerunning, I wanted the fastest attention kernel that actually works on these cards. The RTX PRO 6000 is Blackwell, but prosumer Blackwell: compute capability sm_120a. That distinction turned out to be the whole story. The vllm-omni image ships flashinfer 0.6.14 with four diffusion attention flavors, and I tried them in order of ambition:
| Kernel | Result on sm_120a |
|---|---|
flashinfer cute-dsl (auto-selected on capability ≥ 10) | Dead: uses tcgen05 MMA, which exists only on datacenter sm_100a/103a/110a. Fails fast at job start with expects arch ... but got sm_120a |
flashinfer cutlass | Dead: same tcgen05 requirement; the error message literally says Use backend='fa2' instead |
flashinfer fa2 | Runs, then the ragged-prefill planner int32-overflows on a 357-frame reference’s packed sequence: Trying to create tensor with negative dimension -65994240 |
CUDNN_ATTN | Crashes at 362-frame sequences with a 1080p reference… but see below |
TORCH_SDPA | Works always, ~6-7x slower, yesterday’s 4h21m |
Two things worth underlining. First: every fancy flashinfer path is gated on an instruction my cards physically lack, so “Blackwell” on the box does not mean “datacenter Blackwell” to the kernel selector. Second: the failures all failed fast — thirty seconds at job start instead of four hours in. I’ll take a rude error at minute zero over a polite timeout at hour four every time.
Pinning the sub-kernel, for anyone who needs it: --diffusion-attention-config takes JSON, and the flashinfer backend choice hides inside the quant block, which is only serialized when quantization is “enabled” — so you set both dtypes to bfloat16 (a no-op cast) purely to smuggle flashinfer_backend through:
--diffusion-attention-config '{"default":{"backend":"FLASHINFER_ATTN",
"quant":{"dtype_qk":"bfloat16","dtype_vo":"bfloat16","flashinfer_backend":"fa2"}}}'
(In a systemd Exec=, backslash-escape every quote, or systemd eats them and the container dies on unrecognized arguments. Ask me how I know.)
It earned me nothing here — fa2 is the one that overflows — but the mechanism is undocumented enough that someone will need it.
The three fixes
1. Patch the watchdog with a bind-mount. Extract the file, sed the constant, mount it over the container path in the quadlet. No image rebuild, survives until the next nightly pull:
Volume=/path/to/patched/diffusion_engine.py:/usr/local/lib/python3.12/dist-packages/vllm_omni/diffusion/diffusion_engine.py:ro,Z
30.0 → 3600.0. The post-denoise stage now has an hour instead of thirty seconds.
2. Halve the reference resolution. H3’s output canvas is model-locked — MINIMAX_H3_OUTPUT_SHORT_EDGE = 768, and the pipeline raises if you ask for anything else, so 1344×768 output is non-negotiable. But the reference video is your choice, and every reference frame is VAE-encoded into conditioning tokens that ride along in the packed attention sequence at every single denoise step. The ref went from 1440×1080 to 720×540: four times fewer conditioning pixels, shorter sequence, faster steps, less memory pressure.
3. Back to CUDNN_ATTN. Yesterday’s crash happened at 362 output frames with a 1080p reference. With the 540p reference the packed sequence is dramatically shorter, and a 4-second smoke with the full 357-frame ref (794s inference, peak 72.5 GiB per GPU, clean output) said cudnn now holds. That smoke result is what justified risking the fast kernel on the long job instead of resigning myself to SDPA.
The rerun
Same prompt, same seed, same 362 frames. Watched through py-spy dump --locals on the vLLM-Omni::DiffusionWorker processes (the only progress signal that exists — the API’s progress field stays 0 forever):
- ~2 min/step on cudnn vs ~5.3 min/step on SDPA — a 2.7x speedup, mostly from the kernel, partly from the shorter conditioning sequence
- 49 steps, then VAE decode, then MP4 encode: 1h45m total, versus 4h21m for sampling alone yesterday
- The 30s watchdog, now 3600s, never fired. Yesterday it fired exactly once, at the worst possible moment.
The result — arcade for 7 seconds, a morph around second 8, photorealistic to the end, with the original arcade soundtrack muxed back over it:
I verified the content by extracting frames: t=2s is genuine sprite-art (limited palette, hard edges), t=8s is already fully photographic (chrome, guardrails, ocean haze), t=14s stays cinematic with the same road geometry. The model even kept the blonde passenger.
Epilogue: my automation then lied to me
The finisher script (poll status → download → mux audio → stop the service) downloaded the MP4 perfectly, then declared ERROR: content fetch failed after render completed and stopped the server — its ffprobe duration check raced the curl stream and read a partially-written file. Meanwhile my crash monitor saw the service stop and reported DONE_ENGINE_DOWN: the “engine death” was my own finisher’s systemctl stop. Two layers of automation, two false alarms, one perfectly good 21.9 MB MP4 sitting in /tmp the whole time. Verify artifacts before believing error logs — especially your own.
What I’d tell past me
- Read the timeout constants before you queue a four-hour job.
grep -rn "wait_for\|timeout"in the output path would have found the 30s watchdog in a minute. It cost a full render to learn it existed. - Smoke the failure shape, not the happy shape. A 4s smoke with the full-length reference is what proved cudnn viable; the 4s smokes with short refs that green-lit the original run proved nothing about 362-frame sequences.
- “Blackwell” is two architectures. sm_100/103/110 have tcgen05 MMA; sm_120/121 do not. Kernel selectors that branch on
capability >= 10are wrong for prosumer cards, and the error only appears at runtime. - The input resolution is the speed lever when the output is locked. You can’t negotiate H3’s 1344×768 canvas, but the reference’s pixels are conditioned on at every step. Halve them.
- A finish-line failure is a pipeline bug, not a modeling problem. Nothing about yesterday’s loss was the model’s fault: watchdog, kernel choice, and my own monitoring. All fixable, all fixed, all in this post so the next long render is boring.