STRIDE: Automated Evaluation of Text-to-Trajectory Alignment across Diverse Contexts
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2608. 05745v1 Announce Type: cross Abstract: Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics.
AgentVidBench is a new multi‑hop video question‑answering benchmark designed to evaluate spatial, temporal, and causal reasoning in multimodal large language models (MLLMs). Unlike existing tests that focus on simple scene queries or global summaries, AgentVidBench includes step‑by‑step solution traces to assess whether agents gather the necessary evidence to justify their answers. Experiments with 12 MLLMs show limited single‑turn performance, but integrating these models into agentic workflows improves both accuracy and trajectory scores, establishing AgentVidBench as a comprehensive testbed for future research on agentic video understanding.
The paper introduces Latent Action Driving Annotations (LADA), a three‑stage pipeline that converts large amounts of unlabelled observation‑trajectory data into a language‑conditioned driving model. First, a latent action model with a vector‑quantised bottleneck learns a compact codebook of vehicle intents. Then, a small set of language‑annotated examples trains a vision‑language translator to map observations and instructions into this codebook, and finally a VLA is trained on observation‑latent‑action pairs across the full corpus. Using less than 5% of language annotations, LADA attains a Driving Score of 87.98 and a Success Rate of 70.46% on Bench2Drive, matching or surpassing fully supervised baselines.
arXiv:2606. 18824v1 Announce Type: cross Abstract: Pedestrian trajectory prediction from an ego-centric camera is challenging since it depends on complex interactions with vehicles and scene context, as well as the intention of the pedestrian.
The paper introduces Temporally-Grounded Language Generation (TGLG), a benchmark that tests vision‑language models on their ability to produce semantically accurate and temporally precise utterances in real‑time settings. It identifies perceptual updating and contingency awareness as key capabilities, curates datasets from sports broadcasting and egocentric interactions, and proposes the TRACE metric to jointly evaluate semantic similarity and temporal alignment. The authors also present VLM‑TSI, a model that interleaves visual and linguistic tokens in a time‑synchronized manner, achieving better performance than a strong baseline yet still showing modest overall results, underscoring the challenge of real‑time VLMs.
arXiv:2606. 30266v1 Announce Type: cross Abstract: Motion-language agents must possess the bidirectional capability to both understand human movement (motion-to-text, M2T) and generate it from natural language (text-to-motion, T2M).