Shaduf.Research preview
Real-Time AI Video/LiveAvatar: 45 FPS stream is not a finished conversation

Pool topic: Real-time, continuous, and interactive AI video generation. Question revision: 1. Exact question: What can generate continuous, interactive AI video today, and which setups work at which cost, latency, hardware requirements, and quality?

Checked 9 October 2026 · live-avatar route check

LiveAvatar's 45 FPS does not mean an instant conversation

Authors report: 45.2 generated FPS and 1.21 seconds from audio arrival to first visual output on five H100s, at 720×400 with four-step FP8 inference. Their paper estimates roughly three seconds end to end with transport.

Two different routes

The LiveAvatar paper defines “real-time” as faster-than-playback throughput, not conversational delay. Its 1.21-second first-frame number includes frame-boundary wait, generation and VAE decode, but not a complete voice agent or viewer network. The authors say the estimated three-second delay falls short of seamless two-way interaction. The released repository has a five-GPU streaming inference command, but still lists easy interactive-stream UI and TTS integration as unfinished. Its single-GPU command is labelled offline.

The smaller bill is for a finished clip

An independent deployment benchmark ran the official single-GPU path on one H200: 181.44 seconds from container start to final MP4 for a portrait and 3.63-second narration. It recorded 61,401 MiB peak GPU memory and estimated $0.3184 in Modal compute for that one job. Different hardware, warm state and output boundary make this no comparison to the five-H100 45.2-FPS stream. The dollar figure is not an hourly live-service price.

Before choosing: decide whether the job needs a finished talking clip or an interruptible avatar. For the latter, request a consented, timed audio-to-receiver session with quality acceptance, failures and a complete invoice. None was measured by this pool.

Search published pools, pages, reports, and evidence.