← Back to all news
arXiv Computer Vision August 28, 2026 By Chenyang Wu, Fuchen Long, Binyuan Huang, Xinlong Sun, Xi Chen, Chun-Le Guo, Chongyi Li

Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning

Read the original on arXiv Computer Vision →

The Flow has not summarised this story yet — read it at arXiv Computer Vision.

  • llms
  • rag
  • agents
  • computer-vision
  • multimodal
  • benchmarks
  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

Hugging Face Trending Papers
2d ago

Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning

While generative AI has significantly advanced video editing, existing methods primarily focus on single-shot or short video clips. Editing long videos with multiple instructions remains a formidable...

llmsragagentscomputer-visionmultimodalbenchmarkssafety
More like this →
arXiv Computer Vision
2d ago

FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry

arXiv:2607.18227v2 Announce Type: replace Abstract: In line with the prevailing direction of vision research, we explore the integration of both generation and editing capabilities for video and imag...

By Dingyun Zhang, Lixue Gong, Wei Liu
computer-visionfine-tuningmultimodal
More like this →
arXiv AI
Jul 7

VideoAgent: All-in-One Framework for Video Understanding and Editing

arXiv:2606. 23327v2 Announce Type: replace-cross Abstract: Video editing has become essential in digital media creation, yet existing automated systems are restricted to short segment processing and domain-specific tasks.

By Hengji Zhou, Lingxuan Huang, Jian Wang, Bing Zhou, Si Wu, Lianghao Xia, Chao Huang
llmsagentsmultimodalbenchmarks
More like this →
arXiv AI
4d ago

SEAM: Shot Entity-Attribute Memory for Consistent Short-Drama Generation at Scale

arXiv:2608.22725v1 Announce Type: new Abstract: Short-drama generation has grown into a large, industrialized pipeline, and as it scales from isolated shots to the episode level, visual continuity ha...

By Jiaqi Liu, Maolin Ran, Xiaoyang Lu, Jian Wang, Weiwen Liu, Jianghao Lin, Yong Yu, Weinan Zhang
agentsbenchmarks
More like this →
arXiv AI
Jun 6

Edit-R2: Context-Aware Reinforcement Learning for Multi-Turn Image Editing

arXiv:2606. 05950v1 Announce Type: new Abstract: Text-guided image editing has advanced rapidly with diffusion models and unified multimodal foundation models.

By Yuxiao Ye, Haoran He, Fangyuan Kong, Xintao Wang, Pengfei Wan, Kun Gai, Ling Pan
diffusionreinforcement-learningmultimodalbenchmarks
More like this →
arXiv Computer Vision
3d ago

CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing

arXiv:2608.17566v2 Announce Type: replace Abstract: The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets mainly focus on single editing...

By Fuchen Long, Cong Wang, Zitao Gao, Wenhao Zhong, Yu Cheng, Xiaolu Hou, Yan Li, Xiao Cao, Xinlong Sun, Xi Chen, Yu Liu
benchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea