Paper overview · author-verified

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation

Authors: Chi Zhang, , Xi Qiu, Fangqiu Yi, Xuelong Li

arXiv · 2024

video generationstoryboardcontrollable generationtemporal consistencyscene composition

Publication details

Authors
Chi Zhang, Yuanzhi Liang, Xi Qiu, Fangqiu Yi, and Xuelong Li
Recommended paper citation
Chi Zhang, Yuanzhi Liang, Xi Qiu, Fangqiu Yi, and Xuelong Li. “VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation.” arXiv (2024). arXiv:2412.16677.

Version dates

arXiv first posted
2024-12-21
arXiv last revised
2024-12-21
Source checked
2026-07-31
Verification status
Author-verified on 2026-07-31

Summary

VAST is a two-stage Video-As-Storyboard-from-Text framework: StoryForge converts text into storyboards containing human poses and object layouts, and VisionForge turns those structural plans into videos.

Research paths

How this paper contributes to the site's broader research map.

Research question

Direct text-to-video generation makes it difficult to maintain temporal coherence while controlling subject motion and scene composition.

What the paper contributes

  • Decouples text understanding and video synthesis through an explicit storyboard representation.
  • Uses StoryForge to derive human-pose and object-layout structure from text.
  • Uses VisionForge to generate motion with temporal and spatial coherence from the storyboard.

Evidence and evaluation scope

The paper reports VBench improvements in visual quality and semantic expression over compared methods. The result is tied to the stated benchmark and version; exact scores and model settings belong in citations to the paper's tables.

Scope and limitations

The two-stage design depends on storyboard quality: structural errors from StoryForge can constrain VisionForge. The abstract does not support describing VAST as a universal interface for arbitrary reference images, styles, layouts, and motion controls.

Positioning for related work

VAST belongs to planning-then-generation video methods. Its specific intermediate representation is a storyboard that makes pose and layout explicit before video synthesis.

Related-work context

Zhang et al. introduce VAST, a two-stage text-to-video framework in which StoryForge produces pose- and layout-aware storyboards and VisionForge converts them into temporally and spatially coherent videos.

A concise, neutral description of how this paper can be situated in related work.

Official abstract

Generating high-quality videos from textual descriptions poses challenges in maintaining temporal coherence and control over subject motion. We propose VAST (Video As Storyboard from Text), a two-stage framework to address these challenges and enable high-quality video generation. In the first stage, StoryForge transforms textual descriptions into detailed storyboards, capturing human poses and object layouts to represent the structural essence of the scene. In the second stage, VisionForge generates videos from these storyboards, producing high-quality videos with smooth motion, temporal consistency, and spatial coherence. By decoupling text understanding from video generation, VAST enables precise control over subject dynamics and scene composition. Experiments on the VBench benchmark demonstrate that VAST outperforms existing methods in both visual quality and semantic expression, setting a new standard for dynamic and coherent video generation.

The abstract is reproduced for scholarly identification and remains under the paper publisher/authors’ original copyright; it is not covered by this page’s CC BY license.

Evidence references

What to verifyLocation in the paper
Problem statementAbstract
Method and contributionsAbstract; StoryForge and VisionForge sections
Evaluation statementAbstract; VBench experiments

Primary source: arXiv record (arXiv:2412.16677v1).

How to cite

Cite the paper—not this explainer—for scientific claims. Cite this page only when reusing its original commentary.

Chi Zhang, Yuanzhi Liang, Xi Qiu, Fangqiu Yi, and Xuelong Li. “VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation.” arXiv (2024). arXiv:2412.16677.

Reuse policy

Original explanatory text on this page is licensed under CC BY 4.0 with attribution and a link to this page. Paper title, abstract, figures, and bibliographic metadata are excluded and retain their original rights. Creative Commons Attribution 4.0 International.

Primary sources and resources