Shaduf.
Real-Time AI Video/FastH3 local access, reality checked

Research update 002 · 11 September 2026

FastH3 got local. It did not get live.

FastVideo now publishes local FastH3 Preview paths for Apple Silicon MLX and NVIDIA DGX Spark. That is a real access change. Its own numbers show why “runs locally” is not a stream-speed claim.

The change: FastH3 is no longer a four-B200-only story

Project says: the current FastVideo cookbook supports FastH3 Preview on Apple Silicon MLX after local conversion and on one or two NVIDIA DGX Sparks. The MLX route is T2VA—text-to-video with audio—only. Its FL2VA and Ref2VA paths are not wired there.

That opens a local route for controlled clips and prompt iteration. It does not erase the older high-end FastH3 profile or prove a low-cost 24/7 setup.

The number to keep: 23 seconds of work for one second of preview video

Project says: FastVideo compared the same 832×480, 124-frame, four-step FastH3 clip across an M4 Max, one DGX Spark, and two DGX Sparks. At 24 fps, 124 frames is about 5.17 seconds of video.

Source conditionDecoderReported timeCalculated work per output second
M4 Max, 36 GB unified memoryFull H3 VAE451 s87.3 s
One DGX SparkFull H3 VAE243 s47.0 s
Two DGX SparksFull H3 VAE209 s40.5 s
Two DGX SparksTAEH3 preview119 s23.0 s

The calculation is wall time ÷ (124 ÷ 24). The source calls full H3 VAE its quality path. TAEH3 is a fast preview decoder; it says the shortcut softens fine detail. The fastest reported local result is therefore a prompt-check figure, not a 24 fps feed.

What this changes for a builder

Use local FastH3 when you want a short, reproducible clip and you can accept the hardware and the wait. Keep it in the offline or prebuilt-queue column until a setup proves it can build enough ready footage ahead of playout.

Do not treat a local server process as a live service. The missing measurements are direction-to-ready time, repeated-server behavior, failure rate, ready-buffer depth, actual hardware cost, and what a viewer sees after a miss.

The 16 GB reality check points the same way

Community report: a condition-rich MiniMax H3 workflow record reports 5.2-second, 1024×576 T2V clips in 171.6 seconds with a Turbo option and 379.2 seconds without it on an RTX 5060 Ti 16 GB. It is a useful local baseline lead, not a pool benchmark or a universal 16 GB recommendation.

The honest promise for a 16 GB or local FastH3 route is: “you may make a short clip.” The promise is not: “your viewers will see their next idea now.”

What to inspect before spending

  1. Choose a route: an open Director session, a buffered clip system, or an offline local render.
  2. For a local FastH3 test, keep resolution, frames, steps, seed, decoder, cold time, repeat time, memory, and failures together.
  3. Compute wall seconds per output second. Then decide how much finished footage must wait in the buffer.
  4. Compare full-quality output with preview output before you let a timing chart choose the product.
  5. For a viewer-facing stream, test the late-render fallback and session handoff before calling the system continuous.

Search published pools, pages, reports, and evidence.