Checked 26 September 2026 · vLLM-Omni H3 batching boundary
H3 can schedule between denoise steps. Its four-request test still took longer.
Project says: vLLM-Omni says H3 can admit and retire requests between denoise steps. In its published two-H100 test, four immediate H3 jobs had worse wall time and mean latency when co-batched than when run in ordinary request mode. A scheduler switch is not a live-video result.
Four simultaneous jobs did not become a faster queue
Project says: on two H100 GPUs, four 672×384, 209-frame, 30-step H3 requests submitted at the same time took 174.8 seconds in ordinary request mode, with 111.5 seconds mean latency. With step execution and four co-batched requests, the source reports 182.1 seconds wall time and 175.7 seconds mean latency.
The source explains that H3 denoising is compute-bound: packing several already-large requests does not reduce enough shared compute to offset the extra work. This is one project-maintained benchmark, not a general queue result.
What remains unknown
The comparison used simultaneous arrivals. vLLM-Omni says staggered-arrival admission latency is still unmeasured. It does not establish the time to first video, complete MP4, browser playback, input-to-screen response, quality, retries, cost, or a buffer that stays ahead of viewers.
Do not use the four-request table to predict a 5090, a different clip shape, a FastH3 adapter, or a public stream. Measure those separately.
The two-RTX-5090 route still has a task boundary
Project says: the current two-GPU launch selects the FL2VA task path. The recipe maps T2VA and FL2VA to that path and recommends a 384 GiB host for the offload profile. Its separate validation is one 50-step T2VA job: 124 1344×768 / 24 fps frames in 8 minutes 38 seconds.
That is a capacity result for a selected task path. It does not document a combined all-task service, an FL2VA timing result, or a four-request queue that keeps up.
A server task still needs a usable client
vLLM-Omni’s current H3 API/ComfyUI RFC says its current ComfyUI client accepts one FL2VA frame. The server recipe can represent ordered first-and-last-frame inputs. A direct service task and a chosen client workflow are different checks.
What to test
Start with one pinned task and one request. Then record cold and warm completion, first playable media, host and GPU memory, a second queued request, and what a viewer sees. Only test co-batching with the arrival pattern, clip shape, and latency target you actually need.