No Corners Cut: State-Grounded Transitions for Mid-Stream Prompt Switches in Video Generation
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The paper introduces Statebench, a benchmark for evaluating how well video generators track world states across segments, focusing on past-visible, occluded-process, and complex-transition states. It also proposes Stateagent, a method that maintains an explicit entity-state representation, updates it with new prompts, and uses the resulting state to guide video continuation. Experiments show Stateagent raises the overall state score from 45.2 to 69.3 and improves one‑minute story generation.
arXiv:2609.22283v1 Announce Type: new Abstract: Understanding the design space of streaming video diffusion is essential to exploring its potential for generation quality and computational efficiency...
arXiv:2606. 06991v1 Announce Type: cross Abstract: Online Video Large Language Models (Video-LLMs) have advanced toward seamless human-AI interaction through frame-by-frame processing and proactive responding.
The paper introduces Watch-Think-Interact (WTI), a closed-loop framework for multi-question streaming video reasoning that maintains compact natural-language memory entries linked to video time ranges. WTI decides whether to answer, continue watching, or recall relevant past intervals for each question, avoiding replay of the full history. The authors build a large dataset, WTI-82K, and a training method, Stream-GDPO, achieving state‑of‑the‑art performance on StreamingBench and OVO-Bench.
Zing-0.5 is a 5B autoregressive world model that enables users to explore and influence generated worlds through joint keyboard and online text control. It integrates unified action and text conditioning, event-scale supervision for incremental generation, and low-cost real-time interaction, achieving high scores on WBench Navigation. The authors release model weights, inference code, and a serving implementation to support further research on playable generated worlds.
arXiv:2608.29621v1 Announce Type: cross Abstract: Long-horizon story-driven video generation requires a production agent to coordinate narrative decomposition, state tracking, shot design, prompt con...