Summary
VAST is a two-stage Video-As-Storyboard-from-Text framework: StoryForge converts text into storyboards containing human poses and object layouts, and VisionForge turns those structural plans into videos.
Research paths
How this paper contributes to the site's broader research map.
- Video Generation and World ModelsCore work
VAST inserts an explicit storyboard between text and video, using human pose and object layout as conditioning signals before synthesis.
Research question
Direct text-to-video generation makes it difficult to maintain temporal coherence while controlling subject motion and scene composition.
What the paper contributes
- Decouples text understanding and video synthesis through an explicit storyboard representation.
- Uses StoryForge to derive human-pose and object-layout structure from text.
- Uses VisionForge to generate motion with temporal and spatial coherence from the storyboard.
Evidence and evaluation scope
The paper reports VBench improvements in visual quality and semantic expression over compared methods. The result is tied to the stated benchmark and version; exact scores and model settings belong in citations to the paper's tables.
Scope and limitations
The two-stage design depends on storyboard quality: structural errors from StoryForge can constrain VisionForge. The abstract does not support describing VAST as a universal interface for arbitrary reference images, styles, layouts, and motion controls.
Positioning for related work
VAST belongs to planning-then-generation video methods. Its specific intermediate representation is a storyboard that makes pose and layout explicit before video synthesis.
Related-work context
Zhang et al. introduce VAST, a two-stage text-to-video framework in which StoryForge produces pose- and layout-aware storyboards and VisionForge converts them into temporally and spatially coherent videos.
A concise, neutral description of how this paper can be situated in related work.
Official abstract
Generating high-quality videos from textual descriptions poses challenges in maintaining temporal coherence and control over subject motion. We propose VAST (Video As Storyboard from Text), a two-stage framework to address these challenges and enable high-quality video generation. In the first stage, StoryForge transforms textual descriptions into detailed storyboards, capturing human poses and object layouts to represent the structural essence of the scene. In the second stage, VisionForge generates videos from these storyboards, producing high-quality videos with smooth motion, temporal consistency, and spatial coherence. By decoupling text understanding from video generation, VAST enables precise control over subject dynamics and scene composition. Experiments on the VBench benchmark demonstrate that VAST outperforms existing methods in both visual quality and semantic expression, setting a new standard for dynamic and coherent video generation.
The abstract is reproduced for scholarly identification and remains under the paper publisher/authors’ original copyright; it is not covered by this page’s CC BY license.
Evidence references
| What to verify | Location in the paper |
|---|---|
| Problem statement | Abstract |
| Method and contributions | Abstract; StoryForge and VisionForge sections |
| Evaluation statement | Abstract; VBench experiments |
Primary source: arXiv record (arXiv:2412.16677v1).
How to cite
Cite the paper—not this explainer—for scientific claims. Cite this page only when reusing its original commentary.
Chi Zhang, Yuanzhi Liang, Xi Qiu, Fangqiu Yi, and Xuelong Li. “VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation.” arXiv (2024). arXiv:2412.16677.
Reuse policy
Original explanatory text on this page is licensed under CC BY 4.0 with attribution and a link to this page. Paper title, abstract, figures, and bibliographic metadata are excluded and retain their original rights. Creative Commons Attribution 4.0 International.