论文概要
FreeLong 无需额外训练,在去噪过程中融合全局视频特征的低频部分与局部子序列特征的高频部分,使预训练短视频 diffusion model 同时兼顾长时全局一致性和局部时空细节。
研究问题
直接让短视频 diffusion model 生成更长序列会明显退化;论文将其与空间高频成分下降和时间高频成分上升的频率失真联系起来。
论文贡献
- 分析短视频模型扩展到长序列时的频率分布失真。
- 在去噪时融合低频全局特征和高频局部特征,无需重新训练。
- 支持连贯的长视频与 multi-prompt 生成。
证据与评测范围
论文在多个基础 video diffusion model 上评测 FreeLong,并报告一致性、保真度和多 prompt 转场的改善;准确帧长、底模和指标依实验设置而定。
适用范围与局限
该方法依赖预训练短视频模型本身的能力以及全局/局部特征分解方式。它扩展了时间范围,但并不能自动解决所有长时语义或物理一致性问题。
Related work 定位
FreeLong 是 inference-time、training-free 的长视频生成方法,区别于重训练或蒸馏路线,主要改变去噪阶段的 temporal feature mixing。
Related Work 表述
Lu 等提出 FreeLong,在去噪时融合低频全局特征与高频局部特征,无需训练即可扩展短视频 diffusion model 的长视频生成能力。
这是一段用于说明论文定位的简洁中性表述。
论文官方英文摘要
Video diffusion models have made substantial progress in various video generation applications. However, training models for long video generation tasks require significant computational and data resources, posing a challenge to developing long video diffusion models. This paper investigates a straightforward and training-free approach to extend an existing short video diffusion model (e.g. pre-trained on 16-frame videos) for consistent long video generation (e.g. 128 frames). Our preliminary observation has found that directly applying the short video diffusion model to generate long videos can lead to severe video quality degradation. Further investigation reveals that this degradation is primarily due to the distortion of high-frequency components in long videos, characterized by a decrease in spatial high-frequency components and an increase in temporal high-frequency components. Motivated by this, we propose a novel solution named FreeLong to balance the frequency distribution of long video features during the denoising process. FreeLong blends the low-frequency components of global video features, which encapsulate the entire video sequence, with the high-frequency components of local video features that focus on shorter subsequences of frames. This approach maintains global consistency while incorporating diverse and high-quality spatiotemporal details from local videos, enhancing both the consistency and fidelity of long video generation. We evaluated FreeLong on multiple base video diffusion models and observed significant improvements. Additionally, our method supports coherent multi-prompt generation, ensuring both visual coherence and seamless transitions between scenes.
摘要仅用于学术识别,版权仍归论文作者或出版方所有,不属于本页 CC BY 许可范围。
依据与出处
| 核对内容 | 论文中的位置 |
|---|---|
| 问题陈述 | Abstract |
| 方法与贡献 | Abstract; SpectralBlend Temporal Attention method section |
| 评测结论 | Abstract; experiments on multiple base video diffusion models |
主要核验来源: NeurIPS proceedings record (NeurIPS 2024 proceedings version; arXiv:2407.19918v1 cross-checked).
如何引用
科研结论应引用论文本身;只有在复用本站原创解读时才引用本页。
Yu Lu, Yuanzhi Liang, Linchao Zhu, and Yi Yang. “FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal Attention.” Advances in Neural Information Processing Systems (2024), 37, 131434-131455. https://doi.org/10.52202/079017-4177.
复用许可
本页原创解读采用 CC BY 4.0:复用时须署名并链接本页。论文标题、摘要、图表和书目信息不在此许可范围内,仍保留原有权利。 Creative Commons Attribution 4.0 International.