Research update 005 · 14 September 2026
Four B200s do not clear FastH3’s documented 15-second deadline.
Reactor’s current FastH3 README gives two different high-end numbers: about 15.5 seconds to build a roughly 15-second clip on four B200 GPUs, and about 12.9 seconds on eight. Do not flatten that into “real time.”
One GPU-count change flips the source’s timing
Project says: FastH3’s current documentation uses four B200s as its tested default and says a roughly 15-second clip takes about 15.5 seconds to build there. It also gives about 12.9 seconds on eight B200s. Four is slightly behind the cited clip duration; eight is ahead by the project’s stated numbers.
That does not make eight B200s a proven live-video service. The source does not give this pool an independent queue-drain test, retry rate, rental bill, uptime result, or viewer study. It does give builders one less excuse to hide the hardware count behind a 1× label.
The stream client has a plan for ready clips
Project says: Reactor’s supplied streaming client keeps FastH3 autoplay on, curates the playout front, and uses idle filler while chat is quiet. That can make the start between already-ready clips very short. It does not make a late viewer prompt ready, and it does not remove FastH3’s hard cuts between independently generated scenes.
The demo question has changed: show the GPU count, warm-up, ready seconds, filler policy, and the moment the playout queue drains. Smooth delivery after a clip is ready is useful. It is a different claim from generation keeping up with demand.
The hosted context route just got more expensive
Provider says: fal now lists H3 Max Director at $0.08 per generated second, with a 60-second minimum and public sessions up to 15 minutes. Its stated minimum spend is $4.80; its stated full-session generated-video spend is $72.
Director remains the inspected route that the provider says carries context inside a held WebRTC session. These terms do not prove direction delay, reconnection behavior, moderation, delivery cost, or a clean restart.
Local H3 still buys time, not instant response
Community report: one r/comfyui author reports a roughly 15-second 768×1024 H3 result in about 22 minutes on an RTX 5090 with 64 GB RAM, eight steps, a Turbo sampler, and a LoRA. The post calls the route free of API credits while excluding hardware and electricity.
That is a valuable local-shoot baseline lead. It is not a general 5090 benchmark, an optimized recipe, or a live-stream result. A current ComfyUI development-version issue also records an unresolved high-VRAM interruption report on a 5090 D, which is why version belongs beside hardware.
What to ask before a FastH3 spend
- Which GPU count produced the timing, and which clip length, canvas, prompt, and warm-up plan did it use?
- What is the accepted-input-to-ready distribution, not only the warm median?
- How many finished seconds sit ahead of playout after viewer prompts, retries, and filler?
- What picture and audio appear after the queue drains?
- What does the whole delivered hour cost, including idle hardware and failures?
Until those answers exist, call FastH3 what the source supports: a project-documented, high-end queue route for distinct short scenes.