Shaduf.
Real-Time AI Video/Cost and Throughput

Cost and throughput · checked 21 September 2026

A 5090 FastH3 result is a clip-timing result. Playback still needs a timing result.

Put a clip’s output duration beside time-to-ready and the boundary being measured. Then add queue margin, retries, idle hardware, encoding, delivery, and session resets before calling a route fast or cheap.

FastH3 V2 native canvas: one Radeon report, still far behind playback

Community report: one author reports FastH3 8-Step V2 on a Ryzen AI Max+ 395 / Radeon 8060S system with 128 GB unified memory, Windows, ROCm 10.0.0, and ComfyUI v0.36.0. The stated 768×1344, 243-frame job completed in 33 minutes 9 seconds end to end.

Source-stated conditionReported time boundaryOutput durationDerived work per output secondDecision meaning
FastH3 8-Step V2 · Radeon 8060S · 128 GB unified memory · 768×1344 · 243 frames1,989 s end to endAbout 10.13 s at 24 fpsAbout 196.4 sNative-canvas local render lead, not playback-rate generation.

The arithmetic is 1,989 seconds divided by 243 ÷ 24. The author says VSA did not engage below 12,288 tokens, so a smaller 640×384 case did not gain from FastH3. This is one self-report, not a quality comparison, a transferable AMD result, a cold-start record, or a measure of FL2VA or live capacity.

New 5090 row: complete MP4, still slower than playback

Platform benchmark says: Sogni reports a pinned FastH3 Preview v1 image-to-video test on one RTX 5090 (32 GB). At 24 fps, its 362-frame output is 15.08 seconds long. The stated hot job-start-to-finished-MP4 result was 55 seconds at 480×640 and 126 seconds at 768×1024.

Published conditionReported time boundaryOutput durationDerived work per output secondDecision meaning
FastH3 Preview v1 I2V · 1× RTX 5090 32 GB · 480×640 · 362 frames · 4 steps55 s hot job start to finished MP415.08 s at 24 fpsAbout 3.65 sFaster clip review or prequeue, not playback-rate generation.
Same named route · 768×1024126 s hot job start to finished MP415.08 s at 24 fpsAbout 8.36 sOffline image-anchored clip, not an interaction result.

The arithmetic divides each stated wall time by 15.08 seconds. The source excludes queue and model-load time and says a 480p FastH3 canary after a worker restart took 160 seconds. It reports its own platform test and does not establish a general 5090 baseline, dollar cost, quality parity, retries, first-frame delivery, or stream capacity. This is four-step Preview v1, not the separate eight-step V2 Comfy package.

New server row: full H3 on four B300s

Project says: vLLM-Omni’s current MiniMax H3 recipe reports a warmed FL2VA workload on four B300 GPUs: 209 frames at 1248×768, 50 steps, and 86.964 seconds mean HTTP client latency. At H3’s fixed 24 fps, 209 frames are about 8.71 seconds of output.

Published conditionTime boundaryOutput durationDerived work per output secondDecision meaning
Full H3 FL2VA · 4× B300 · 1248×768 · 209 frames · 50 steps86.964 s mean HTTP client latencyAbout 8.71 s at 24 fpsAbout 9.99 sRemote offline/API workflow. Not a playback-rate, queue, or interaction result.

The calculation is 86.964 seconds divided by 209 ÷ 24 fps. The source excluded one warmup and averaged three requests. It does not include a rental price, cold start, endpoint queue, client rendering, transport, retries, quality review, or a viewer’s wait.

Do not compare the H3 server row with FastH3 by name alone

The vLLM route is a full-H3, 50-step FL2VA condition. Its current integration RFC says FastH3 requires a selected adapter fused at server start, and it is not switchable through the remote ComfyUI LoRA control. A few-step FastH3 result needs its own task, artifact, attention path, quality, hardware, warm state, and time boundary before it belongs in the same table.

Consumer channel: source-stated 22.1 fps, named conditions

Community project says: its FastH3 Live v1.2.0 route sustains 22.1 fps on one RTX 5090 (32 GB) under Windows 11 with ComfyUI. Its named setup uses a dense four-step conversion, a 448×448 canvas, a prequeue, retimed playout, and a 721-scene library.

Speed labelWhat the source saysDecision meaningMissing evidence
24 fpsFastH3 authors its 362 frames for roughly 15.08 seconds of native-speed video.This is the motion rate that needs to be matched for normal-speed playback.One-5090, 448×448 repeated performance at 24 fps.
22.1 fpsThe community project’s current source-stated sustained pace on its named machine and workflow.Closer to native motion, yet still a sub-24-fps local-channel condition.Independent timing, quality, uptime, delivery, and cost result.
18 fpsThe author’s earlier example plays 362 frames for 20.11 seconds.75% native motion buys more playout time for the next clip.Audience acceptance for people and other mid-speed action.

FastH3 local: published timing, not a live claim

Project says: FastVideo compared the same 832×480, 124-frame, four-step FastH3 clip at 24 fps on an M4 Max, one DGX Spark, and two DGX Sparks. A 124-frame clip is about 5.17 seconds of output. The table converts its stated times into a decision number.

Published local conditionDecoderReported timeApprox. work per output secondWhat it is for
M4 Max with 36 GB unified memoryFull H3 VAE451 s87.3 sLocal clip, not playout
One DGX SparkFull H3 VAE243 s47.0 sLocal clip, not playout
Two DGX SparksFull H3 VAE209 s40.5 sLocal clip, not playout
Two DGX SparksTAEH3 preview119 s23.0 sFast prompt check; source reports softer fine detail

Current priced route: H3 Max Director

Provider says: the current fal page lists $0.08 per generated second, a 60-second minimum, and public sessions up to 15 minutes. That makes the stated minimum $4.80 and a stated full 15-minute session $72 in generated-video spend.

Unknown: actual billed time, direction-to-change latency, delivery cost, moderation cost, and whether an audience sees a clean restart. This arithmetic is not an operating quote.

Use this record

API response timeTime to the service’s returned artifact or status. It is not automatically time to a playable, delivered, or visible result.
Native-motion fpsThe rate the source frames were authored to represent. A lower delivered fps can extend playout time while slowing motion.
Generated-minute costProvider bill or compute rate divided by usable generated seconds.
Delivered-hour costGenerated work plus idle time, retries, buffering, encoding, transport, and sessions that fail before a viewer sees value.
Ready secondsFinished footage waiting ahead of the viewer. This is the margin that protects a queue-based stream.

Search published pools, pages, reports, and evidence.