FairSSL is a multimodal self‑supervised learning framework that treats data heterogeneity as a fairness resource instead of a limitation. It replaces strict alignment with a subject‑aware Variance‑Invariance‑Covariance Regularization objective, enforcing alignment only across segments from the same subject. The method includes segment‑based pooling for variable‑length modalities and regularizes representations to promote within‑subject variability, cross‑modal and cross‑subject invariance, and decorrelation, theoretically bounding the score gap between protected groups and empirically outperforming baselines on heterogeneous multimodal datasets.
By Jiaee Cheong, Abtin Mogharabin, Paul Liang, Hatice Gunes, Sinan Kalkan
UrbanVLA is a Vision‑Language‑Action framework designed to enable delivery robots to navigate large‑scale urban environments using long‑horizon route instructions. The model aligns noisy route waypoints with visual observations and plans trajectories, trained through a two‑stage pipeline of supervised fine‑tuning on simulated data and reinforcement fine‑tuning on mixed simulation and real‑world data. Experiments show UrbanVLA outperforms strong baselines by over 55% on the SocialNav task and demonstrates reliable real‑world navigation in large urban settings.
By Anqi Li, Zhiyong Wang, Jiazhao Zhang, Minghan Li, Yunpeng Qi, Zhibo Chen, Zhizheng Zhang, He Wang
The study investigates whether vision–language models (VLMs) can infer human engagement from gameplay videos using the GameVibe Few‑Shot dataset across nine first‑person shooter games. Three VLMs were tested under six prompting strategies—including zero‑shot, theory‑guided prompts based on Flow, GameFlow, Self‑Determination Theory, and MDA, and retrieval‑augmented prompting—evaluating both pointwise engagement prediction and pairwise prediction of engagement change. Results show that zero‑shot predictions are weak and often do not beat simple majority‑class baselines; retrieval‑augmented prompting improves pointwise prediction in some cases, while pairwise prediction remains difficult, and theory‑guided prompts do not reliably help and may reinforce superficial shortcuts.
"whyItMatters":"The findings highlight a perception–understanding gap in current VLMs, indicating that while they can recognize visible gameplay cues, they still struggle to robustly infer human engagement across games."
By Ziyi Wang, Qizan Guo, Rishitosh Singh, Xiyang Hu
Rewind-IL is a training‑free online safeguard for generative action‑chunked imitation learning policies. It uses a zero‑shot failure detector based on Temporal Inter‑chunk Discrepancy Estimate (TIDE) and a state‑respawning mechanism that returns the robot to a verified safe intermediate state. The system builds a checkpoint library offline with a vision‑language model and monitors self‑consistency online, rewinding execution to the latest safe checkpoint when a failure is detected, thereby improving reliability in long‑horizon manipulation tasks.
By Gehan Zheng, Sanjay Seenivasan, Matthew Johnson-Roberson, Weiming Zhi
LEARN-TS is a method for multivariate time‑series anomaly detection that uses a frozen language model to generate two semantic representations: one that captures window‑specific observation context without exact numerical values, and another that provides a fixed, dataset‑agnostic normality prompt. These representations guide a patch‑masked reconstruction process, allowing the model to estimate normality discrepancy and produce timestamp‑level anomaly evidence. Experiments on four benchmarks show that LEARN‑TS achieves the best mean performance in 13 of 16 dataset‑metric comparisons, and ablation studies confirm the benefits of observation conditioning, joint normality alignment, and semantic references.
By Jahyeob Koo, Kio Yun, Byoungmo Koo, Jun-Geol Baek
Diffusion-2BC is a hybrid training method that combines a diffusion denoising objective with an auxiliary deterministic behavior‑cloning loss over a shared visual encoder. The auxiliary loss is used only during training, while inference remains diffusion‑based. Experiments on the Claw environment and CARLA navigation show that Diffusion‑2BC reduces mean mask‑distance error by about 10% compared to a diffusion baseline and by 85% compared to standard deterministic behavior cloning, and it enables the agent to travel farther and exhibit multimodal route choices.
By Bruno Maciel Machado, Eric Aislan Antonelo
The paper investigates a specific failure mode in vision‑language‑action models, termed instruction‑action binding, where models respond to language and vision separately but fail to combine them to select the correct action under counterfactual changes. Through behavioral analyses of fine‑tuned policies, the authors show that failed rollouts often preserve source behavior or switch to other demonstrated tasks, indicating that language is not ignored but mis‑bound. They propose Equivariant Counterfactual Training (ECT), which supplies counterfactual demonstrations and a paired loss to enforce correct action selection, achieving significant performance gains across simulated and real‑world benchmarks.
By Hung-Jen Chen, Yu-Hsun Hou, Yan-Hong Chen, Yan-Fu Chen, Binghua Cai, Min Sun, Chun-Yi Lee
The paper introduces MedCertAIn, a framework that uses data‑driven priors to enhance uncertainty estimation in multimodal clinical models. By integrating cross‑modal similarity and modality‑specific corruptions into neural network priors, the authors improve predictive performance and selective prediction for in‑hospital mortality risk using MIMIC‑IV and MIMIC‑CXR data. The results demonstrate competitive accuracy and notable gains over deterministic and stochastic baselines, underscoring the potential of such priors for reliable, uncertainty‑aware clinical decision support.
By L. Juli\'an Lechuga L\'opez, Tim G. J. Rudner, Farah E. Shamout
GPEC is a lightweight pre‑LLM Gaussian Process Embedding Correction that refines visual representations for cardiac ultrasound caption generation. By inserting a residual correction layer between the visual projection and the language model, it aligns projected embeddings with annotation‑guided targets derived from structured video annotations. Experiments show that GPEC improves caption similarity and content alignment while adding less than 0.05 s of inference overhead, all without fine‑tuning the pretrained multimodal backbone.
By Arefeh Rezaei
The paper introduces RateAR, a benchmark of AR images and videos that captures perceptual factors such as object placement, scale, and shadow consistency. Eleven commercial vision‑language models (VLMs) are evaluated on this benchmark, with VLM‑based quality predictions showing strong correlation with human judgments (Spearman’s ρ up to 0.8695). Using contextual prompting, the authors build an automated AR content adjustment system that, in a 21‑participant study, improved placement and size coherence of virtual content for over 90% of users.
By Elias Rotondo (Duke University), Lin Duan (Duke University), Yanming Xiu (Duke University), Sangjun Eom (Duke University), Conrad Li (Duke University), Maria Gorlatova (Duke University)
The paper introduces Align Then Reason (ATR), a multilingual lip‑sync judge that aligns frame‑level lip representations with phonetic units of a candidate text line and then uses a language model to evaluate both content and timing. ATR achieves significant improvements over existing baselines on a seven‑language benchmark, with mean AUC gains of up to 59.4% for 2B reasoners and similar gains across other LLM families. The method also transfers well to unseen languages and outperforms lip‑reading baselines on real dubbing tasks such as dub‑line reranking and script‑to‑clip assignment.
By Rui Liu, Bhavin Jawade, Haoqi Li, Shivam Mehta, Karan Saxena, Yinghong Lan, Cameron R. Wolfe
Watch Your Speech (WYS) is a video‑to‑speech synthesis framework that uses textual conditioning to resolve the one‑to‑many mapping problem inherent in silent talking‑face videos. It fuses textual context with video sequences via an attention‑based embedding module and employs a conditional flow matching objective to produce high‑fidelity, phonemically accurate speech. Experiments on LRS2 and LRS3 show that WYS sets new state‑of‑the‑art results in audio‑visual synchronization while maintaining competitive word‑error rates, and subjective tests confirm near‑human naturalness.
By Gunwoo Lee, Yoori Oh, Yoseob Han
FutureWorlds is a framework that learns robotic world models by generating alternative future predictions and using them as learning signals. It combines candidate construction, history maintenance, and learning from relative quality through a multimodal discrete autoregressive model and diverse beam search. The MemSPO algorithm further optimizes the world model by converting video trajectory rewards into group-relative advantages, leading to significant improvements in LPIPS scores across RT-1, BridgeV2, and RoboCasa datasets.
By Hao Wu, Shengju Qian, Weiyan Wang, Fan Xu, Fan Zhang, Yuanpeng He, Qingsong Wen, Yuxuan Liang
Dyna3 is a training‑free framework that extends the depth foundation model DA3 to perform 4D dynamic scene reconstruction without fine‑tuning. By leveraging DA3’s cross‑view features and a best‑match search, it distinguishes static surfaces from moving objects, and uses vision‑language models to generate semantic prompts for SAM 3 to achieve precise instance‑level segmentation. Experiments on four datasets show Dyna3 outperforms correspondence‑trained methods, improving dynamic object segmentation by +5.5 pp, speeding pose estimation 13×, and reducing memory usage 4–8×.
By Xinhao Xiang, Weiyang Li, Zhijie Zheng, Abhijeet Rastogi, Jiawei Zhang
The paper introduces MOSAIC, a method for personalized federated vision‑language models that addresses domain heterogeneity by focusing on class‑specific cross‑domain residuals. It constructs a decision‑aware harmfulness score to identify residuals that hurt image‑text decision margins, then uses a low‑rank residual adapter with shared class factors and private domain factors, an image‑conditioned gate, and harmful‑pair‑aware reweighting to refine local updates. Experiments on Office31, OfficeHome, and DomainNet100 show consistent improvements in macro‑client top‑1 accuracy across various domain‑shift scenarios.
By Wentao Yue, Qingyu Mao, Tianyou Lai, Ahmed M. Abdelmoniem, Qilei Li
UnifiedAttack introduces a benchmark for testing the safety of large multimodal models (LMMs) in tasks that combine text and image to produce harmful content. The benchmark focuses on the additional harm that arises from cross‑modal synergy and includes filtered multimodal samples and synthesized disinformation queries. A synergistic hijacking framework—comprising In‑Context Reskinning (ICR) and Cognitive Planning Injection (CPI)—is proposed to expose vulnerabilities, and extensive evaluations show that UnifiedAttack consistently bypasses current alignment defenses in state‑of‑the‑art architectures.
By Bingjun Luo, Jialin Guo, Tony Wang, Siqi Li
The paper introduces Adaptive Tversky Policy Optimization (ATPO), a reinforcement learning framework for multi‑label video safety detection. ATPO uses an Adaptive Tversky Reward (ATR) that dynamically adjusts false‑positive and false‑negative penalties, allowing controllable precision‑recall trade‑offs. Experiments on SafeWatch‑Bench and XD‑Violence demonstrate significant performance gains, raising the Jaccard Index from 40.66 to 75.44 on SafeWatch‑Bench‑Real.
By Guangyu Yang, Jingbiao Mei, Mingsheng Sun, Jinghong Chen, Yingtong Bu, Pengda Qin, Da Chen, Bill Byrne
FocusGraph is a framework for selecting keyframes in egocentric long‑video question answering. It uses a lightweight Scene‑Graph LLM Selector to identify query‑relevant clips from graph‑based captions, then extracts keyframes with Patch‑wise Sparse‑Flow Retention (PSFR) before feeding them to a multimodal large language model for answer generation. The method achieves state‑of‑the‑art performance on FindingDory and HourVideo while reducing question‑time inference cost.
By Tatiana Zemskova, Solomon Andryushenko, Ilya Obrubov, Viktoriia Khoruzhaia, Ekaterina Eroshenko, Ekaterina Derevyanka, Dmitry Yudin
DynamicVLA is a latency‑aware Vision‑Language‑Action model designed for dynamic object manipulation, featuring a compact 0.4B architecture and a convolutional vision encoder for efficient multimodal inference. It employs a continuous inference schedule that overlaps reasoning and execution, and a Latent‑aware Action Streaming mechanism that discards stale action prefixes to maintain action‑time alignment. The authors also introduce the Dynamic Object Manipulation (DOM) benchmark, comprising 200K synthetic episodes and 2K real‑world episodes, and demonstrate that DynamicVLA improves dynamic manipulation success in simulation and on real robots.
By Haozhe Xie, Beichen Wen, Jiarui Zheng, Zhaoxi Chen, Fangzhou Hong, Haiwen Diao, Ziwei Liu
VIDA (Visually-Dependent Ambiguity) is a new dataset comprising 2,500 curated instances designed to test whether multimodal machine translation models can use visual input to resolve ambiguous source spans. The authors also introduce Disambiguation-Centric Metrics, which employ an LLM-as-a-judge classifier to verify correct span-level resolution. Experiments with recent large vision‑language models show that visual disambiguation remains difficult, and that chain‑of‑thought supervised fine‑tuning yields stronger out‑of‑distribution performance, especially for collective‑noun ambiguities.
By Jingheng Pan, Xintong Wang, Longyue Wang, Liang Ding, Weihua Luo, Chris Biemann