Shaduf.
Real-Time AI Video/FastH3 V2 scope recheck

Research update 011 · 20 September 2026

FastH3 V2 has an official transition template. Its own model card says that transition was never distilled.

The decision: do not promise a first/last-frame FastH3 V2 transition yet. FastVideo’s current model card says its 8-Step V2 checkpoint is text-to-audio-video only. ComfyUI’s current image-to-video template says the same checkpoint accepts first and last frames. That is a source conflict, not a workflow result.

Two official paths, one incompatible claim

Project says: FastVideo’s current 35B FastH3 8-Step V2 card describes an eight-forward checkpoint with 80% Video Sparse Attention. It says FL2VA and Ref2VA were not distilled, and names four B200 GPUs for its trained default path.

ComfyUI says: its current image-to-video template points to the same FastH3 V2 file and accepts optional first_frame and last_frame inputs. The template says no images means T2VA, while connected frames mean FL2VA; it excludes Ref2VA.

Neither statement proves a render. A template can expose inputs without settling what the distilled checkpoint preserves. A model card can state a training boundary without proving every framework integration is impossible. Until a pinned run resolves that difference, a frame-anchored transition is Unknown.

Why this kills the easy continuity story

First/last-frame conditioning is the difference between trying to bridge two chosen images and generating a fresh text-only clip. If you need a character, shot, or ending frame to carry into the next scene, do not replace that requirement with an eight-step headline.

The V2 card also does not provide a consumer-GPU timing, VRAM requirement, price, retry rate, quality score, time to first media, or stream result. Four B200 GPUs and eight transformer forwards do not answer what will happen on your workstation or in front of a viewer.

The five-minute compatibility card

  1. Pin the FastVideo and ComfyUI revisions, exact checkpoint, task, scheduler, VSA/attention path, GPU, and host RAM.
  2. Run one short T2VA control, then the same duration with the supplied first and last frames.
  3. Record cold and warm time, errors, output length, and whether the specified boundary images are visibly anchored.
  4. If the run changes checkpoint or task, fails, or does not honour the frame boundaries, do not market it as FastH3 V2 FL2VA.

That test settles scope. It does not settle speed, quality, queue depth, cost, or live delivery; measure those separately.

What changed for a builder

FastH3 V2 is still worth inspecting as a faster text-to-audio-video route. It is not yet a dependable local continuity route. If a transition matters, start with the task contract before you download a 35B checkpoint or build a playout system around it.

Search published pools, pages, reports, and evidence.