arXiv AI

SGA: Plug&Play Geometric Verification for Educational Video Synthesis

arXiv:2607. 18116v1 Announce Type: new Abstract: Recent work leverages Large Language Models (LLMs) to generate executable code for pedagogical animations using libraries such as Manim.

arXiv AI
Aug 11

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

arXiv:2607. 27380v2 Announce Type: replace-cross Abstract: Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt.

By Haodong Li, Tianfei Ren, Xiaoxiao Ma, Chunmei Qing, Zhen Fang, Sipeng He, Ziyu Guo, Haoyu Wu, Juanxi Tian, Yihang Zou, Ruichuan An, Dongzhi Jiang, Boxue Yang, Ji Xie, Xu Huang, Wenhao Yan, Jialv Zou, Zhengrong Yue, Yaxin Luo, Xiaotong Li, Yuzhu Wang, Junyan Ye, Jinjing Zhao, Zehui Chen, Lin Chen, Renye Yan, Feng Zhao, Pheng-Ann Heng
arXiv Computer Vision
Aug 25

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents

LangDriveCTRL is a natural‑language‑controllable framework that edits real‑world driving videos by representing each video as an explicit 3D scene graph, separating a static background from dynamic object nodes. It employs a feedback‑driven agentic pipeline where an Orchestrator translates user instructions into executable graphs that coordinate specialized multi‑modal agents—Object Grounding, Behavior Editing, and Behavior Reviewer—to align text with scene nodes, generate and refine multi‑object trajectories, and ensure photorealism through a video diffusion tool and Video Reviewer. The system supports object node editing (removal, insertion, replacement) and multi‑object behavior editing, achieving nearly twice the instruction alignment of prior state‑of‑the‑art methods while preserving photorealism, structural integrity, and traffic realism.

By Yun He, Francesco Pittaluga, Ziyu Jiang, Matthias Zwicker, Manmohan Chandraker, Zaid Tasneem
arXiv AI
Jun 2

Beyond End-to-End Video Models: An LLM-Based Multi-Agent System for Educational Video Generation

arXiv:2602. 11790v2 Announce Type: replace Abstract: Although recent end-to-end video generation models demonstrate impressive performance in visually oriented content creation, they remain limited in scenarios that require strict logical rigor and precise knowledge representation, such as instructional and educational media.

By Lingyong Yan, Jiulong Wu, Dong Xie, Weixian Shi, Deguo Xia, Jizhou Huang
Hugging Face Trending Papers
Jul 2

SimWorlds: A Multi-Agent System for Dynamic 3D Scene Creation

LLM agents are increasingly used to translate natural language into 3D scenes in a procedural way, but existing systems focus on static output. Dynamic 4D scenes from text alone, in which liquids flow, particles emit, rigid bodies cascade, and articulated mechanisms move, remain largely unexplored despite their value as editable content and as physics-grounded training data for video generation and embodied AI.

arXiv AI
4d ago

Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes

Code4Scene is a benchmark that evaluates coding agents on constructing and editing 3D scenes in Unreal Engine. It tests agents on two tasks: construction, where they must build a scene from open‑ended language, and editing, where they must recover a target scene from reference images while preserving everything else. The benchmark measures task fulfillment, artifact integrity, and physical validity, revealing that construction and editing performance are correlated but not interchangeable, with agents struggling most with spatial composition and precise edits.

By Xiaokang Ye, Siddhant Hitesh Mantri, Zimeng Chen, Edward Zhang, Zhaoxu Zheng, Yuanheng Li, Yizhao Chen, Tianyang Huang, Lianhui Qin
arXiv AI
Aug 21

ChronoAgentic: A Code-based Multi-Agent World Simulator for Physically Grounded Simulation Construction

arXiv:2605. 14398v3 Announce Type: replace Abstract: Video-based world models generate visually plausible rollouts, but since they infer dynamics in latent states, they enforce no explicit physical constraints: contacts drift, shapes distort, and motion loses consistency.

By Hongyu Wang, Jingquan Wang, Ashvin Anilkumar, Bocheng Zou, Radu Serban, Dan Negrut
arXiv AI
Sep 25

PPTBench: Can Coding Agents Reconstruct the Visual World through Structured, Editable Slides

PPTBench is a new benchmark that tests coding agents’ ability to reconstruct scientific flow diagrams from arXiv papers into editable PowerPoint slides. The dataset contains 500 tasks, each requiring agents to produce a single PPTX page with native, editable objects, and a four‑stage Agentic Judge evaluates validity, semantic correctness, rendering quality, and fine‑grained visual quality. Across 31 model configurations, the best score is 67.80, with a median of 19.47, showing that while agents can generate valid PPTX files, they still struggle with semantic and visual accuracy, especially text details.

By Xiaoqiu Wang, Yizhe Chi, Wenyi Li, Deyao Hong, Zhihan Shan, Mingju Gao, Kaisen Yang, Youjie Zheng, Calvin Xiao, Qinhuai Na