论文概要
VAST 是两阶段的 Video-As-Storyboard-from-Text 框架:StoryForge 将文本变成人体姿态与物体布局明确的 storyboard,VisionForge 再把这一结构规划生成视频。
研究问题
直接 text-to-video 很难同时维持时间一致性,并精确控制主体运动和场景构图。
论文贡献
- 通过显式 storyboard 表示将文本理解与视频合成解耦。
- StoryForge 从文本中生成包含人体姿态和物体布局的结构规划。
- VisionForge 根据 storyboard 生成具备时间和空间一致性的运动视频。
证据与评测范围
论文报告在 VBench 上相对比较方法改善视觉质量和语义表达;该结论与特定 benchmark 和论文版本绑定,准确分数与模型设置应引用实验表格。
适用范围与局限
两阶段设计依赖 storyboard 的质量,StoryForge 的结构错误会限制 VisionForge。摘要并不支持把 VAST 描述为可统一处理任意参考图、风格、布局和动作控制的通用接口。
Related work 定位
VAST 属于先规划、再生成的视频方法,其特定中间表示是显式编码 pose 与 layout 的 storyboard。
Related Work 表述
Zhang 等提出 VAST:StoryForge 先生成包含姿态与布局的 storyboard,VisionForge 再将其转换为时间和空间一致的视频。
这是一段用于说明论文定位的简洁中性表述。
论文官方英文摘要
Generating high-quality videos from textual descriptions poses challenges in maintaining temporal coherence and control over subject motion. We propose VAST (Video As Storyboard from Text), a two-stage framework to address these challenges and enable high-quality video generation. In the first stage, StoryForge transforms textual descriptions into detailed storyboards, capturing human poses and object layouts to represent the structural essence of the scene. In the second stage, VisionForge generates videos from these storyboards, producing high-quality videos with smooth motion, temporal consistency, and spatial coherence. By decoupling text understanding from video generation, VAST enables precise control over subject dynamics and scene composition. Experiments on the VBench benchmark demonstrate that VAST outperforms existing methods in both visual quality and semantic expression, setting a new standard for dynamic and coherent video generation.
摘要仅用于学术识别,版权仍归论文作者或出版方所有,不属于本页 CC BY 许可范围。
依据与出处
| 核对内容 | 论文中的位置 |
|---|---|
| 问题陈述 | Abstract |
| 方法与贡献 | Abstract; StoryForge and VisionForge sections |
| 评测结论 | Abstract; VBench experiments |
主要核验来源: arXiv record (arXiv:2412.16677v1).
如何引用
科研结论应引用论文本身;只有在复用本站原创解读时才引用本页。
Chi Zhang, Yuanzhi Liang, Xi Qiu, Fangqiu Yi, and Xuelong Li. “VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation.” arXiv (2024). arXiv:2412.16677.
复用许可
本页原创解读采用 CC BY 4.0:复用时须署名并链接本页。论文标题、摘要、图表和书目信息不在此许可范围内,仍保留原有权利。 Creative Commons Attribution 4.0 International.