The paper introduces an evidence‑guided detector‑localizer‑reasoner system for text‑centric image forensics, addressing the challenges posed by AI‑generated content. It combines an image‑level authenticity detector, a localizer that extracts tampered regions, and an MLLM‑based reasoner that generates structured forensic reports grounded in the detected evidence. The system employs iterative difficulty‑aware mining and report‑mask consistency post‑processing, achieving a score of 0.638 and ranking second in the ACM Multimedia 2026 GenText‑Forensics Challenge.
By Peifeng Liu, Bin Li, Qingsong Zhang, Yangxin Yu, Leqing Chen, Xiaoye Qiu
The paper introduces a new framework for estimating full-degree-of-freedom egomotion directly from asynchronous optical flow captured by event cameras. By decoupling the differential epipolar constraint into angular and linear components and applying a first-order approximation, the authors derive a polynomial formulation that yields the first algebraic minimal 5‑point solver for this problem. An accelerated solver that truncates high‑order angular velocity terms is also proposed, enabling real‑time performance in high‑speed scenarios, and extensive tests show superior accuracy and robustness compared to traditional synchronous methods.
By Shuo Pan, Banglei Guan, Bin Li, Zhenbao Yu, Zibin Liu, Zi Wang, Yang Shang, Qifeng Yu
The paper surveys the transition from textual chain-of-thought reasoning to action-grounded reasoning in autonomous driving, highlighting that driving decisions require continuous actions that mirror the spatiotemporal structure of the physical world. It reviews 171 papers, categorizing 130 methods into four main types—language-based, visual-spatial, latent-dynamic, and externalized reasoning—along with 13 subtypes linked to specific regions of interest. The authors argue that the future of driving agent reasoning lies in intermediate representations that are grounded in reality, linked to real-time actions, and verifiable within safety-critical systems.
By Zhengxu Tang, Xiaozhou Zhang, Guofeng Cui, Ziyu Gong, Zi Wang, Yunfei Shi, Ruifeng Deng, Chengzhi Qi, Ke Chen, Sachin Patil, Tianjun Xiao, Langechuan Liu, Pichao Wang
The paper introduces OpenRef, a benchmark for Referring Expression Comprehension (REC) designed for open‑world scenarios. OpenRef expands beyond simple settings by including diverse visual domains, variable target counts (multi‑target and none‑target), and a rich vocabulary with proper nouns, polysemous words, and ordinal terms. It also proposes new evaluation metrics—F1 for grounding accuracy and N3R for negative expression rejection—and presents a training‑free Multi‑task Consistency Checker (MCC) that improves model performance with a single click.
By Zongjian Wu, Lei Zhang
The paper introduces a generalizable deformation learning framework that reconstructs 3D objects by deforming a category-level shape template to match a monocular observation. It employs a geometry-guided feature modeling mechanism to enrich foundation features with template topology, creating a geometry-aware representation that is explicitly correlated with the target observation for precise deformation. A view-adaptive feature aggregation module further bridges the gap between the fixed template and arbitrary target views by leveraging multi-view template features and camera poses, ensuring robust feature alignment across diverse viewpoints.
By Yiyao Ma, Kai Chen, Zhongxiang Zhou, Zhuheng Song, Dongsheng Xie, Zelong Tan, Rong Xiong, Qi Dou
The paper presents a hardware‑accelerated instance segmentation framework tailored for resource‑constrained lunar robotics, addressing low‑light perception, limited compute, and radiation‑induced hardware faults. It introduces Activation Variance Informative Sampling (AVIS), a label‑free calibration method that selects samples based on activation variance, and deploys a YOLO‑based model on a Deep Learning Processor Unit with architectural tweaks to reduce CPU fallback and ensure bounded latency. A software‑level criticality analysis estimates fault exposure, guiding mitigation that reduces global criticality by 31.7%, while AVIS with bias correction recovers 69.8% of quantization‑induced accuracy loss at 309 ms latency and 5.7 W power consumption.
By Siddhant Shete, Hilmi Dogu K\"uc\"uker, Udo Frese, Frank Kirchner
VIPS is a benchmark for vehicle‑to‑infrastructure cooperative autonomous driving that uses pseudo‑simulation to combine vehicle and infrastructure observations, enabling scalable yet realistic evaluation of robustness and error propagation without full simulation. The paper also introduces CoS‑V2X, a cooperative planning framework that employs sparse representations to model vehicle‑infrastructure interactions efficiently and robustly under heterogeneous observations.
By Hoonhee Cho, Jae-Young Kang, Giwon Lee, Hyemin Yang, Heejun Park, Kuk-Jin Yoon
MultiGraspNet is a multitask 3D vision model that simultaneously predicts feasible poses for both parallel and vacuum grippers, allowing a single robot to handle multiple end effectors. Trained on the aligned GraspNet-1Billion and SuctionNet-1Billion datasets, it generates graspability masks that quantify the suitability of each scene point for successful grasps. With only 15.75 M parameters, the model achieves fast inference on a single GPU and demonstrates competitive performance against single-task models while reducing computational cost, as shown in extensive experiments and real‑world tests on a single‑arm multi‑gripper setup.
By Stephany Ortuno-Chanelo, Paolo Rabino, Enrico Civitelli, Tatiana Tommasi, Raffaello Camoriano
Sim2Signal is a benchmark designed to systematically measure the Sim-to-Real gap in traffic signal control by decomposing it into observation, action, transition, and reward gaps. The study evaluates 18 mitigation methods across 33 gap settings and 10 calibrated networks from five real-world locations, finding that direct transfer degrades performance but mitigation effectiveness varies by network and gap type. The most effective approaches tend to estimate the specific changes caused by each gap rather than relying on domain randomization or invariant representations.
By Ferdous Al Rafi, Susrik Mukherjee, Latika Liladhar Dekate, Jennifer Yawa Lavoe, Huaiyuan Yao, Shlok Mohanty, Longchao Da, Xuesong Zhou, Hua Wei
GEM is a Generative LiDAR world model that uses a deformable Mamba architecture to better handle the disorder of LiDAR point clouds and distinguish dynamic objects from static structures. The model tokenizes LiDAR sweeps, unsupervisedly disentangles dynamic and static features, and applies a tri‑path deformable Mamba for selective scanning and adaptive gating fusion, improving spatial‑temporal understanding. Experiments show GEM outperforms existing methods across multiple benchmarks, and it can be paired with a planner and BEV controller for autonomous rollout and "what‑if" scenario generation.
By Yang Wu, Zhaojiang Liu, Qiang Meng, Youquan Liu, Renliang Weng, Jianjun Qian, Jian Yang, Jin Xie
DailyBench is a unified benchmark designed to evaluate AI-generated image detectors on both modern full-image synthesis and object-level manipulation. It comprises two subsets: FakeBench, featuring high‑quality images from recent open‑source and commercial generative models, and ManipulationBench, containing subtle local edits applied to real images using advanced image‑conditional models. Experiments show that detectors with high accuracy on older datasets perform poorly on DailyBench, revealing significant robustness gaps.
By Xin Jiang, Hao Tang, Junyao Gao, Meiqi Cao, Fei Shen, Dongming Zhang, Yongdong Zhang
DiDrive introduces a risk‑aware hierarchical diffusion framework for offline reinforcement learning in autonomous driving. It combines a low‑level risk‑gated encoder with a high‑level contextual modulator to filter redundant state information, and a 3DICE policy optimization that reduces out‑of‑distribution overestimation and stabilizes gradients. On the CARLA benchmark, DiDrive outperforms baselines such as IQL, CQL, and Diffusion‑QL, achieving an 85% success rate and a 4295.68 average reward in dense traffic with 60 vehicles.
By Qisong Guo, Jingtang Chen, Zhilin Chen, Pei Xu, Mingjian Fu, Wenxi Liu, Yuanlong Yu
The paper presents a deep denoising autoencoder (DAE) approach for non‑invasive detection of blood flow in arteriovenous fistulas (AVFs) using waveform data processed by a one‑level discrete wavelet transform. The DAE performs dimensionality reduction and reconstruction, producing a latent representation that achieves 93% accuracy in detecting AVF dysfunction and 92% accuracy in identifying patient‑specific characteristics. Lightweight versions of the model can maintain performance on less powerful devices, indicating practical applicability for routine AVF monitoring.
By Li-Chin Chen, Yi-Heng Lin, Li-Ning Peng, Feng-Ming Wang, Yu-Hsin Chen, Po-Hsun Huang, Shang-Feng Yang, Yu Tsao
ShallowStream is a framework for streaming video understanding that uses the shallow layers of a multimodal large language model (MLLM) to encode frames and build a lightweight index. During streaming, it maintains an always‑on index via the KV cache of shallow layers, and at query time it scores context frames using shallow‑layer attention and selects diverse evidence for answering. The approach matches the performance of leading streaming methods while cutting per‑frame prefill latency and 10‑second end‑to‑end latency by up to 52.1× and 11.9×, respectively.
By Jitai Hao, Ke Yang, Qiang Huang, Jun Yu
Make‑It‑Poseable is a feed‑forward framework that treats 3D character posing as a skinning‑free latent‑space transformation. It decouples shape deformation from fixed mesh connectivity, using a latent posing transformer, dense pose representation, and an adaptive completion module with bipartite‑matched latent loss. Experiments show it outperforms existing baselines, generalizes to varied morphologies, and supports 3D authoring tasks such as part replacement and refinement.
By Zhiyang Guo, Ori Zhang, Jax Xiang, Alan Zhao, Zhenxun Yuan, Wengang Zhou, Houqiang Li
TAPVid-MV is a new benchmark for tracking any point in 3D across multiple synchronized camera views. It comprises 284 sequences, 1,142 calibrated camera streams, and 109,769 point tracks, covering indoor and outdoor domains and derived from various modalities such as depth, LiDAR, SLAM, and simulation. The dataset is visually verified, and evaluation shows that current multi‑view trackers do not consistently outperform monocular trackers, highlighting geometry recovery as a key bottleneck.
By Skanda Koppula, Frano Rajic, Abdullah Faiz Ur Rahman, Yi Yang, Ignacio Rocco, Jeet Thakwani, Rishabh Kabra, Andrew Zisserman, Joao Carreira, Siyu Tang, Carl Doersch, Gabriel Brostow
A new method named CW‑Net converts the reasoning of an autonomous vehicle’s AI into understandable concepts, allowing humans to see why the car behaves a certain way. By translating the AI’s internal logic into clear explanations, the system helps users predict when a self‑driving car might make mistakes. This approach bridges the gap between complex machine learning processes and human comprehension.
By Adam Zewe | MIT News
The paper proposes a foundational ontology to represent contradictions in dialogue-based human‑robot interactions. Using METHONTOLOGY and Activity Theory, it defines dialogues and contradictions through natural language, set theory, and first‑order logic, and introduces three new principles for human‑robot dialogue. The work aims to create a formal, interoperable framework applicable across HRI and human‑agent interaction domains.
arXiv:2609.00369v1 Announce Type: new
Abstract: Generating co-speech gestures that are temporally coherent, semantically aligned with speech, and grounded with surrounding objects remains challenging...
By Vida Adeli, Soroush Mehraban, Jacob Rommann, Harrison Sanborn, Cole Clifford, Babak Taati
arXiv:2602.16365v2 Announce Type: replace-cross
Abstract: Flexible endoscopic continuum manipulators offer high dexterity and access to complex anatomy, but nonlinear hysteresis limits feedforward co...
By Junhyun Park, Chunggil An, Myeongbo Park, Ihsan Ullah, Sihyeong Park, Minho Hwang