Pool topic: Real-time, continuous, and interactive AI video generation. Question revision: 1. Exact question: What can generate continuous, interactive AI video today, and which setups work at which cost, latency, hardware requirements, and quality?
Checked 2 October 2026 · TaoLiveAIGC project report
TaoMate-H3 publishes a playable first chunk in 17.287 seconds on eight H20s
Project says: its three-step MiniMax H3 runtime generates audio-video in chunks. The cited first-playable benchmark is for a ten-second, 480×864 T2AV job on eight H20 96 GB GPUs. It is not a measured live viewer response.
Three clocks, three answers
| Project-reported measure | Time | Boundary |
|---|---|---|
| Pure DiT for ten seconds of output | 14.810 s | Excludes load, text encoding, VAE decode and media encoding. |
| First final chunk latent | 6.148 s | Not playable media. |
| Benchmark first playable video | 17.287 s | Includes video VAE decode and H.264 publication; excludes internal audio preparation. |
The released streaming runtime is T2AV. Its documented command writes a final MP4. FL2AV is listed as planned before 15 October; Ref2AV has no release date. The four-GPU option is documented, but the published timing uses eight H20s.
The practical choice
Do not treat the downloadable three-step LoRA as the complete streaming system. A separate first-person RTX 5090 ComfyUI report used a converted adapter for offline clips; that author still reported about 14 seconds before one avatar response. It did not validate the official clean-KV runtime on a 5090.
For a show, request one timestamped run from cold start through the first playable audio-video at a receiver, several later blocks, one prompt change, a missed block and the bill. The project has not published that record, so sustained playback, viewer latency, accepted quality and delivered-hour cost remain Unknown.