ViMax: Agentic Video Generation
arXiv:2606. 07649v1 Announce Type: cross Abstract: Long-form video generation requires systematic narrative planning and visual consistency that current short-clip methods cannot provide.
arXiv:2607. 24756v1 Announce Type: cross Abstract: What gets lost when memory becomes media?
arXiv:2606. 07649v1 Announce Type: cross Abstract: Long-form video generation requires systematic narrative planning and visual consistency that current short-clip methods cannot provide.
Short-drama generation has grown into a large, industrialized pipeline, and as it scales from isolated shots to the episode level, visual continuity has become a critical bottleneck. Current agent fra...
arXiv:2608.22725v1 Announce Type: new Abstract: Short-drama generation has grown into a large, industrialized pipeline, and as it scales from isolated shots to the episode level, visual continuity ha...
arXiv:2610.08102v1 Announce Type: new Abstract: Conversational MLLM agents are increasingly expected to assist in professional workflows, from AI research and engineering design to product management...
arXiv:2605. 16716v5 Announce Type: replace-cross Abstract: Text-to-video (T2V) generation has rapidly progressed in visual fidelity, yet its ability to faithfully represent multiple cultures within a single prompt remains underexplored.
arXiv:2605. 16716v4 Announce Type: replace-cross Abstract: Text-to-video (T2V) generation has rapidly progressed in visual fidelity, yet its ability to faithfully represent multiple cultures within a single prompt remains underexplored.
arXiv:2610.00097v1 Announce Type: new Abstract: Recent diffusion and autoregressive models have substantially improved text-to-video generation, yet producing coherent long-form story videos with con...
As generative multimedia evolves from static image synthesis to complex, interleaved visual narratives, a foundational bottleneck has emerged: the judgment crisis. While human perception naturally synthesizes the temporal and logical flow of a story, automated evaluation systems remain largely "blind" to sequential continuity, often failing to distinguish between a coherent narrative and a semantically shuffled or contradictory sequence.
arXiv:2607. 07051v1 Announce Type: cross Abstract: Conversational image editing requires preserving not only visible content, but also content that temporarily disappears across turns.
arXiv:2608.29621v1 Announce Type: cross Abstract: Long-horizon story-driven video generation requires a production agent to coordinate narrative decomposition, state tracking, shot design, prompt con...
arXiv:2604. 25220v2 Announce Type: replace Abstract: Data videos combine animated visualizations with synchronized narration to communicate quantitative information and are widely used in journalism, education, and public communication.
MIRAGE is a controlled study that examines how multimodal personal agents use historical evidence when conversation state changes. The study keeps evidence, questions, and scoring constant while varying only the conversation state, then checks if agents can determine answerability, recover the correct source, and answer from it. Results across seven multimodal backbones show distinct failure regimes before and after compaction, heavy reliance on context continuity by open-weight models, and mixed effects of retrieval pressure on source attribution.