Omni2Web: Benchmarking Audiovisual Website Development
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2608. 02673v1 Announce Type: cross Abstract: Speech editing for content creation requires precise control over both what an edit should do and where it should apply.
OmniEdit-Bench introduces a comprehensive benchmark for instruction-based video editing (IVE), addressing limitations of existing datasets by covering spatial, temporal, audio, and reference-based editing tasks and distinguishing explicit from implicit instructions. The evaluation framework assesses editing quality across accuracy, preservation, realism, and consistency, using human judgments and vision-language models, and incorporates an accuracy-aware penalty to ensure instruction fidelity. Experiments reveal that current IVE models perform poorly, highlighting the need for improved methods.
arXiv:2609.08275v1 Announce Type: new Abstract: Recent multi-shot audio-video generators can produce increasingly coherent and cinematic outputs, but coherence does not imply the ability to execute e...
arXiv:2608.16344v3 Announce Type: replace Abstract: Indic quality estimation (QE) and automatic post-editing (APE) data is spread across separate releases, so no single resource supports training and...
arXiv:2606. 08415v1 Announce Type: cross Abstract: While recent text-guided video editing models excel at elementary tasks (e.
AVENUE is a new benchmark and evaluation framework for audio‑video editing that includes 1,291 source clips and 7,957 editing instructions covering audio‑targeted, video‑targeted, and coupled edits. It introduces a sample‑specific, modality‑aware evaluation that specifies the intended change and the content that must remain unchanged. The study applies this framework to joint, sequential, and separate editing models, revealing that existing models often alter unintended modalities, highlighting a key challenge in controllable AV editing.