A Dive into Text-to-Video Models
Related stories
Training Design for Text-to-Image Models: Lessons from Ablations
Video generation models as world simulators
We explore large-scale training of generative models on video data. Specifically, we train text-conditional diffusion models jointly on videos and images of variable durations, resolutions and aspect ratios.
M4V: Multimodal Mamba for Efficient Text-to-Video Generation
arXiv:2506. 10915v2 Announce Type: replace-cross Abstract: Text-to-video generation has significantly enriched content creation and holds the potential to evolve into powerful world simulators.
OSVE: One Step Video Editing with One Step Diffusion Models
arXiv:2607. 19895v1 Announce Type: cross Abstract: Text-guided video editing with diffusion models is impractically slow, hindered by costly multi-step sampling and inversion.
A Systematic Evaluation of Positional Bias in Multi-Video Summarization with MLLMs
Multimodal Large Language Models (MLLMs) are increasingly used for video understanding, yet their reliability under multi-video inputs remains poorly understood. We study positional bias in multi-video summarization, where the quality of a per-video summary can change with the video's input slot even when the underlying content is unchanged.
MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models
arXiv:2606. 06853v1 Announce Type: cross Abstract: The new era has witnessed a remarkable capability to extend Vision-Language Models (VLMs) for tackling tasks of video understanding.
BLM-SGAN: Bidirectional Language Modeling for Semantic-Spatial Text-to-Image Generation
arXiv:2606. 08847v1 Announce Type: cross Abstract: Despite the success of image generation from text descriptions, it still faces challenges that are difficult to overcome in domains such as natural language processing (NLP) and computer vision (CV).
Inference-Time Concept Suppression and Video-Centric Evaluation for Text-to-Video Models
arXiv:2607. 14194v1 Announce Type: cross Abstract: Text-to-video (T2V) generators can synthesize realistic and temporally coherent videos, but controllably removing a target concept from a generator remains difficult.
Motion Attribution for Video Generation
arXiv:2601. 08828v2 Announce Type: replace-cross Abstract: Despite the rapid progress of video generation models, the role of data in influencing motion is poorly understood.
Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data
arXiv:2607. 09876v1 Announce Type: cross Abstract: Automatically retrieving videos from large camera-trap datasets remains challenging.
AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report
arXiv:2607. 18367v1 Announce Type: new Abstract: Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly.