Decision guide · checked 22 September 2026
Choose the control and deployment contracts, then test the claim
A model name does not tell you whether you are buying a group-choice loop, an interactive session, a remote GPU service, a clip queue, an unattended local channel, or a short local render.
Eight routes, eight jobs
| Need | Route to inspect | Evidence status | Constraint that changes the answer |
|---|---|---|---|
| A remote H3 service for a client or team | vLLM-Omni H3 server | Project says: one service can expose T2VA, FL2VA, and Ref2VA to an OpenAI-compatible video API. | Task, server revision, FastH3 adapter at startup, GPU and host memory, cold/warm time, client compatibility, queue, and delivery. The published full-H3 four-B300 row is not live. |
| A shared audience should select the next story beat | Yoroll H3 Superfast / YoLive | Provider says: proposals and votes guide the next scene; 10 seconds at 768p/24 fps with native audio generate in four seconds on eight B200 GPUs. | Vote window, moderation, input-to-screen delay, public uptime, cost, and what survives into the next scene. |
| Viewer input should alter a held stream | fal H3 Max Director | Provider says: persistent WebRTC tracks, carried context, and mid-session text directions. | Public session limit, current $0.08/second billing, reconnect, and moderation design. |
| A frame-carried FastH3 creative loop | FastVideo Dreamverse FastH3 profile | Project says: four visible GPUs by default, 124-frame 768×1344 audio-video segments, and last-frame conditioning into the next segment. | Cold readiness, warm segment time, conditioning fidelity, ready-buffer margin, recovery, cost, and the visible response to a direction. |
| A chat-shaped feed of short scenes | FastH3 queue and playout | Project says: separate build and play queues can feed paced output; its cited 15-second build is 15.5 s on four B200s and 12.9 s on eight. | GPU profile, warm-up, ready-buffer depth, autoplay/filler policy, hard cuts, and what viewers see after a miss. |
| An unattended local AI-TV channel | One-RTX-5090 FastH3 community route | Community project says: its named dense FastH3 workflow sustains 22.1 fps at 448×448. | Retimed motion, source variant, attention path, prequeue, scene independence, host memory, image quality, and repeatability. |
| A local FastH3 clip or possibly a frame-anchored transition | FastH3 8-Step V2 ComfyUI template | Project and ComfyUI say: a new eight-forward package and template exist. | Scope conflict: FastVideo says T2AV only; the template says T2VA and FL2VA, not Ref2VA. No local timing, quality, or hardware result was found. |
| One controlled local shot | MiniMax H3 through a local framework such as ComfyUI | Provider and community reports say: local base clips and an optimized dynamic-VRAM path exist; one 5090 report gives a condition-rich offline result. | Task mode, system RAM, model pack, resolution, steps, software version, and wait tolerance. |
Pick remote serving when
You need the creator UI and the H3 machine to be different systems, or you want several clients to reuse one service. vLLM-Omni’s current H3 recipe records an OpenAI-compatible route. Its stated full-H3 four-B300 FL2VA baseline takes 86.964 seconds to return 209 frames—about 8.71 seconds at 24 fps. Build an offline or queue-ahead workflow first.
Pick a group-choice loop when
You want participation to create anticipation for the next beat instead of a viewer controlling the current one. Yoroll’s launch announcement says YoLive uses proposals and votes. The timing headline is promising; do not choose the route until you know the selection rule, moderation delay, response time, delivery behavior, and cost.
Pick Dreamverse when
You need a published FastH3 application design that tries to carry a visual boundary forward rather than merely play distinct ready clips. The current Dreamverse README calls for four visible GPUs by default. Treat it as a four-GPU evaluation route: time the cold readiness state, then inspect a warm handoff before you write a continuity promise.
Pick Director when
You need the product shape of an open session. The current provider page says it streams through WebRTC and carries context into later directions. Plan for a finite session from the first sketch, and price $4.80 for even the stated 60-second minimum.
Baseline card before you spend a weekend
- Write who can propose, who selects, what changes, and whether the active scene can change.
- State client host, inference host, task, adapter load point, server revision, GPU, host memory, and any carried conditioning frame.
- Record cold, warm, accepted-input, ready, HTTP response, delivered-start, and visible-change timestamps.
- For a channel or queue, record ready clips, playback speed, filler, retries, and black or silent fallback.
- Label the result as provider documentation, project report, community report, or a measured test. Do not blend them.