Checked 29 September 2026 · vLLM-Omni H3 recipe
Four H3 denoising steps still took 11.7 seconds to return a 4.4-second clip
Project says: on eight NVIDIA B300 GPUs, its FastH3 Dense T2VA setup returned a complete 1344×768, 4.4-second MP4 in 11.7 and 11.8 seconds after one excluded warmup. The faster model step did not make this short request finish before its playback clock.
The measured boundary
Under the recipe's named eight-GPU setup, four-step FastH3 spent 2.36–2.37 seconds in the diffusion engine. The complete response took about five times that. The same source gives 25.8 and 26.4 seconds for base 50-step H3 on this test. That is about 2.2× faster for the complete response, not the roughly 6.9× denoiser-only gain.
This short-clip test and the separately reported eight-B300, 10.125-second FastH3 clip returned in 8.678–8.710 seconds are different profiles. Do not use either number as a general H3 speed or viewer-visible live result.
Before selecting an adapter
The current recipe also lists different Turbo LoRAs for text/first-frame and reference-video tasks. It validates the exact artifact against the task, step count and flow shift. A four-step label alone is not a setup. Choose the task and file first, then test a complete clip at the duration and quality you need.
Unknown: this pool has not run the setup, inspected the video, measured a queue, first playable frame, delivered stream, repeated-session reliability, or cost. The recipe is maintained on a changing main branch; pin a commit before reproducing it.