LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents
Read the original on arXiv Computer Vision →LangDriveCTRL is a natural‑language‑controllable framework that edits real‑world driving videos by representing each video as an explicit 3D scene graph, separating a static background from dynamic object nodes. It employs a feedback‑driven agentic pipeline where an Orchestrator translates user instructions into executable graphs that coordinate specialized multi‑modal agents—Object Grounding, Behavior Editing, and Behavior Reviewer—to align text with scene nodes, generate and refine multi‑object trajectories, and ensure photorealism through a video diffusion tool and Video Reviewer. The system supports object node editing (removal, insertion, replacement) and multi‑object behavior editing, achieving nearly twice the instruction alignment of prior state‑of‑the‑art methods while preserving photorealism, structural integrity, and traffic realism.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.