Shaduf.
Real-Time AI Video/Eight B300s clear FastH3’s complete-MP4 playback threshold

Checked 23 September 2026 · vLLM-Omni FastH3 timing check

Eight B300s can finish a FastH3 clip before playback ends. This does not establish a live stream.

The update: vLLM-Omni’s maintainer report says a dedicated FastH3 T2VA service returned a complete 10.125-second MP4 in 8.678 and 8.710 seconds on eight B300 GPUs. This is the first examined route here with a documented complete-MP4 result below its own playback duration.

What the number measures

Project says: the stated profile uses FastH3 Dense/Data-Free T2VA, 1344×768 output at 24 fps, five sigma positions and four transformer forwards. Its timer starts when a synchronous request is submitted and stops when the complete MP4 returns.

The 243-frame output carries 10.125 seconds of video and audio. The two stated result times are 8.678 and 8.710 seconds. That is about 1.16× the output duration. The same report lists 5.175-second outputs in 4.602 and 4.396 seconds, and 15.083-second outputs in 14.177 and 14.059 seconds.

The hardware and task are the headline

This is not a desktop or general H3 result. The published setup uses eight NVIDIA B300 GPUs, one dedicated FastH3 replica, eight-way encoder and Ulysses parallelism, tiled VAE decode, and a pinned artifact. It is T2VA only. The report describes FastH3 as a load-time-fused dedicated student, not a request-switchable first/last-frame transition service.

The older vLLM-Omni full-H3 FL2VA condition remains a different route: four B300 GPUs returned about 8.71 seconds of output in 86.964 seconds. Do not average those rows or call one a speedup over the other.

What it does not prove

  • The report defines “real time” as a complete MP4 arriving before its playback ends. It does not measure streaming delivery or time to first frame.
  • Startup, compilation, and an excluded warmup sit outside the stated timing interval.
  • It does not establish queue behavior, concurrent capacity, browser decode, visible response, retries, uptime, cost, or a frame-carried continuation.
  • The report says its raw benchmark bundle is pending publication. Treat the figures as maintainer-published results with a stated method, not an independently audited benchmark.

The honest next test

Before calling this a live route, record cold boot to ready, request acceptance, complete MP4, playable media, first visible frame, queue depth, a second request, and a failed or late request. Then price the full system. A clip can beat its own playback clock and still leave a viewer waiting.

Search published pools, pages, reports, and evidence.