Shaduf.Research preview
Real-Time AI Video/H3 long-video output is not a live H3 stream

Checked 28 September 2026 · vLLM-Omni H3 v0.30.0

H3 can make a longer clip in vLLM-Omni. It is not a live stream.

Project says: vLLM-Omni v0.30.0 adds an opt-in Ref2VA route beyond H3's normal 15-second request. It generates bounded audio-video windows, carries the previous window's latent tail, and returns a completed clip. The project reports one 30.667-second output on eight B300 GPUs. It does not report the full wait.

What the release supports

The reported example used original MiniMax H3 Ref2VA weights, not a FastH3 adapter: eight NVIDIA B300 GPUs, 1344×768, 24 fps, 50 steps, three windows, and a 736-frame final file. The source describes video and audio overlap, optional prompts for each window, and an option to condition generation on an input soundtrack. The resulting audio is reconstructed, not an exact copy of that soundtrack.

The 300-second continuation limit in the recipe is a request guard, not a verified five-minute output. Continuation currently needs Ref2VA request execution with uncached denoising; it does not support step execution or latent-mask editing. The accumulated latent still grows as the clip gets longer.

What to test before calling it continuous

The report does not provide a complete wall time, cost, repeated-run rate, first playable frame, or a measured join. It does not show an interactive viewer session. If you need a longer finished clip, compare a short full-generation request with continuation on the same model and hardware, then inspect video and audio at the join. If you need a live channel, measure the viewer-visible output and response to a new direction instead.

Search published pools, pages, reports, and evidence.