Pool topic: Real-time, continuous, and interactive AI video generation. Question revision: 1. Exact question: What can generate continuous, interactive AI video today, and which setups work at which cost, latency, hardware requirements, and quality?
Checked 8 October 2026 · research-to-client check
ShotStream reports 16 FPS on one H200. Its public example saves a file.
Paper reports: 15.95 generated FPS at 832×480 on a single NVIDIA H200 for a multi-shot story model. The published example is not a live viewer session.
What the number measures
ShotStream uses a Wan2.1-T2V-1.3B-based next-shot model, not MiniMax H3. Its authors describe changing prompts between shots and report 15.95 generated FPS on one H200. The output in the reference entrypoint is saved at 16 FPS. That is close to playback rate, with no published margin for encoding, transport or retries. The paper was released in March 2026; this is a newly checked route for this pool, not a new October launch. Read the paper and test conditions.
What a reader can run from the repo
The listed inference command reads prepared captions from a CSV. Its entrypoint writes an MP4 after inference returns. The reviewed example supplies neither a live prompt client nor a receiver stream. That does not rule out a working interactive model; it limits what this public release demonstrates.
Next test: on the published checkpoint, timestamp a changed shot prompt, first encoded frame, first receiver-visible frame and one failure/retry. Check whether accepted output holds 16-FPS playout over multiple shots, then price the full H200 session. Those measurements are not published in the inspected sources.