Research update 002 · 11 September 2026
FastH3 got local. It did not get live.
FastVideo now publishes local FastH3 Preview paths for Apple Silicon MLX and NVIDIA DGX Spark. That is a real access change. Its own numbers show why “runs locally” is not a stream-speed claim.
The change: FastH3 is no longer a four-B200-only story
Project says: the current FastVideo cookbook supports FastH3 Preview on Apple Silicon MLX after local conversion and on one or two NVIDIA DGX Sparks. The MLX route is T2VA—text-to-video with audio—only. Its FL2VA and Ref2VA paths are not wired there.
That opens a local route for controlled clips and prompt iteration. It does not erase the older high-end FastH3 profile or prove a low-cost 24/7 setup.
The number to keep: 23 seconds of work for one second of preview video
Project says: FastVideo compared the same 832×480, 124-frame, four-step FastH3 clip across an M4 Max, one DGX Spark, and two DGX Sparks. At 24 fps, 124 frames is about 5.17 seconds of video.
| Source condition | Decoder | Reported time | Calculated work per output second |
|---|---|---|---|
| M4 Max, 36 GB unified memory | Full H3 VAE | 451 s | 87.3 s |
| One DGX Spark | Full H3 VAE | 243 s | 47.0 s |
| Two DGX Sparks | Full H3 VAE | 209 s | 40.5 s |
| Two DGX Sparks | TAEH3 preview | 119 s | 23.0 s |
The calculation is wall time ÷ (124 ÷ 24). The source calls full H3 VAE its quality path. TAEH3 is a fast preview decoder; it says the shortcut softens fine detail. The fastest reported local result is therefore a prompt-check figure, not a 24 fps feed.
What this changes for a builder
Use local FastH3 when you want a short, reproducible clip and you can accept the hardware and the wait. Keep it in the offline or prebuilt-queue column until a setup proves it can build enough ready footage ahead of playout.
Do not treat a local server process as a live service. The missing measurements are direction-to-ready time, repeated-server behavior, failure rate, ready-buffer depth, actual hardware cost, and what a viewer sees after a miss.
The 16 GB reality check points the same way
Community report: a condition-rich MiniMax H3 workflow record reports 5.2-second, 1024×576 T2V clips in 171.6 seconds with a Turbo option and 379.2 seconds without it on an RTX 5060 Ti 16 GB. It is a useful local baseline lead, not a pool benchmark or a universal 16 GB recommendation.
The honest promise for a 16 GB or local FastH3 route is: “you may make a short clip.” The promise is not: “your viewers will see their next idea now.”
What to inspect before spending
- Choose a route: an open Director session, a buffered clip system, or an offline local render.
- For a local FastH3 test, keep resolution, frames, steps, seed, decoder, cold time, repeat time, memory, and failures together.
- Compute wall seconds per output second. Then decide how much finished footage must wait in the buffer.
- Compare full-quality output with preview output before you let a timing chart choose the product.
- For a viewer-facing stream, test the late-render fallback and session handoff before calling the system continuous.