Pool topic: Real-time, continuous, and interactive AI video generation. Question revision: 1. Exact question: What can generate continuous, interactive AI video today, and which setups work at which cost, latency, hardware requirements, and quality?
Checked 10 October 2026 · text-driven avatar route
Hallo-Live reports 20 FPS. Its public command still saves a file.
Authors report: 20.38 generated frames per second and 0.94-second model latency for text-to-audio-video on two H200 GPUs. The published inference entry point writes a final MP4. The released code saves video at 24 fps, above the paper’s 20.38 generated FPS. This comparison does not establish sustained 24-fps playout or a timed browser conversation.
What changed for a builder
The April paper measures a text-driven model that generates a talking character and speech together. Its authors compare it with an Ovi teacher under their test conditions. This is not the same input job as LiveAvatar's audio-driven portrait, so their FPS and first-output numbers are not a head-to-head ranking.
The public repository now provides inference code and checkpoints. Its documented command saves videos in an output folder. A two-GPU option overlaps VAE decoding with generation, but the code restricts that option to text-to-video, not image-conditioned inference. A streaming model design is not yet a documented WebRTC or RTMP client. If the paper throughput applies to the saved-video configuration, it is about 85% of 24-fps playback; queue margin and visible gaps need a session test.
Do not plan around a 24 GB card yet
In a maintainer reply, the default inference setup peaks around 60 GB GPU memory even with automatic T5 offload. The maintainer describes 24 GB as a theoretical target requiring FP8 and stronger offload, not a released measured configuration. Other issue authors ask about custom images, audio and video inputs, and report lip or chunk-boundary defects. These are individual reports, not a measured failure rate.
First useful test: pick one fixed text prompt, record cold load, first decoded frame, whole MP4 time, VRAM and lip/audio acceptance. Only then build a receiver path and time prompt-to-visible change, interruptions and a complete two-H200 bill. The pool has not run this test.