Shaduf.
Real-Time AI Video/Four B200s do not clear FastH3’s documented 15-second deadline

Research update 005 · 14 September 2026

Four B200s do not clear FastH3’s documented 15-second deadline.

Reactor’s current FastH3 README gives two different high-end numbers: about 15.5 seconds to build a roughly 15-second clip on four B200 GPUs, and about 12.9 seconds on eight. Do not flatten that into “real time.”

One GPU-count change flips the source’s timing

Project says: FastH3’s current documentation uses four B200s as its tested default and says a roughly 15-second clip takes about 15.5 seconds to build there. It also gives about 12.9 seconds on eight B200s. Four is slightly behind the cited clip duration; eight is ahead by the project’s stated numbers.

That does not make eight B200s a proven live-video service. The source does not give this pool an independent queue-drain test, retry rate, rental bill, uptime result, or viewer study. It does give builders one less excuse to hide the hardware count behind a 1× label.

The stream client has a plan for ready clips

Project says: Reactor’s supplied streaming client keeps FastH3 autoplay on, curates the playout front, and uses idle filler while chat is quiet. That can make the start between already-ready clips very short. It does not make a late viewer prompt ready, and it does not remove FastH3’s hard cuts between independently generated scenes.

The demo question has changed: show the GPU count, warm-up, ready seconds, filler policy, and the moment the playout queue drains. Smooth delivery after a clip is ready is useful. It is a different claim from generation keeping up with demand.

The hosted context route just got more expensive

Provider says: fal now lists H3 Max Director at $0.08 per generated second, with a 60-second minimum and public sessions up to 15 minutes. Its stated minimum spend is $4.80; its stated full-session generated-video spend is $72.

Director remains the inspected route that the provider says carries context inside a held WebRTC session. These terms do not prove direction delay, reconnection behavior, moderation, delivery cost, or a clean restart.

Local H3 still buys time, not instant response

Community report: one r/comfyui author reports a roughly 15-second 768×1024 H3 result in about 22 minutes on an RTX 5090 with 64 GB RAM, eight steps, a Turbo sampler, and a LoRA. The post calls the route free of API credits while excluding hardware and electricity.

That is a valuable local-shoot baseline lead. It is not a general 5090 benchmark, an optimized recipe, or a live-stream result. A current ComfyUI development-version issue also records an unresolved high-VRAM interruption report on a 5090 D, which is why version belongs beside hardware.

What to ask before a FastH3 spend

  1. Which GPU count produced the timing, and which clip length, canvas, prompt, and warm-up plan did it use?
  2. What is the accepted-input-to-ready distribution, not only the warm median?
  3. How many finished seconds sit ahead of playout after viewer prompts, retries, and filler?
  4. What picture and audio appear after the queue drains?
  5. What does the whole delivered hour cost, including idle hardware and failures?

Until those answers exist, call FastH3 what the source supports: a project-documented, high-end queue route for distinct short scenes.

Search published pools, pages, reports, and evidence.