The paper introduces ReImaGin, a method that uses image generation models as a flexible visual reasoning tool for multimodal large language models. Unlike traditional fixed-function vision tools, ReImaGin accepts natural language commands and can perform open-ended visual operations such as removing occlusions or creating floorplans from multiple views. Experiments on six diverse visual reasoning tasks show that ReImaGin outperforms both text-only reasoning and specialist vision-tool baselines, achieving up to a 25% improvement.
By Nishad Singhi, Hector Garcia Rodriguez, Aditya Arora, Marcus Rohrbach, Anna Rohrbach
NeuroSymbEAD is a large‑scale neuro‑symbolic caption dataset that builds an ego‑centric knowledge graph of static and dynamic objects on the KITTI‑360 dataset, annotating classes, categories, heading directions, orientations, and distances from the ego‑vehicle. The dataset generates multilevel textual captions that serve as a lightweight representation of an ego‑centric scene map, enabling outdoor scene‑map reconstruction, visual recognition, and object grounding. Baselines for driving common sense and traffic/scene understanding are established, and the dataset is benchmarked using pre‑trained grounding and learned auto‑regressive captioning networks to support vision‑language and foundation models for traffic‑scene explanation, 3D reasoning, and interpretable autonomous‑driving perception.
By Muhammad Ahmed Ullah Khan, Mohammed Elamine, Sheikh Talha Uddin, Didier Stricker, Sk Aziz Ali, Muhammad Zeshan Afzal
GeoLAM is a framework that learns geometry‑grounded latent actions from unlabeled human videos. It uses future‑frame reconstruction with a frozen geometric feature hierarchy and motion supervision from a 4D geometry teacher to capture 3D displacement, image‑plane motion, and surface‑orientation changes. After pretraining, the representation serves as transition targets for a world‑action model trained on robot demonstrations, enabling denoised latent actions and executable action chunks without requiring hand‑pose annotations or future‑video generation during deployment.
By Yifan Xie, Hekun Tian, Jinkun Liu, YuAn Wang, Qiao Sun, Wenbo Ding
arXiv:2609.16637v1 Announce Type: cross
Abstract: Efficient perception is central to robotic systems operating under constrained computation, memory, and latency budgets. Knowledge transfer from larg...
By Yanick C. Tchenko, Felix Mohr, Hicham Hadj-Abdelkader, Hedi Tabia
Intrinsic Robot Rewarding (IRR) leverages existing vision‑language‑action (VLA) systems to evaluate a robot’s own outcomes and provide feedback for policy improvement. By using successful demonstration endpoints as task‑specific references and the policy’s frozen visual encoder as the feature space, IRR adds a reference bank and scoring operation to the current pipeline without requiring a separate evaluator or additional perception backbone. The approach aims to reduce integration effort, reward computation cost, and human outcome scoring while enabling learning from the data already available in industrial robot systems.
By Tobias Schaffer, Mohab Elkhayat, Daniela Nicklas, Mustafa Almohamad, Elham Al-Fuqara
The paper introduces the Neverwhere Visual Parkour Benchmark Suite, a collection of over sixty hyper‑photo‑realistic 3D Gaussian Splatting reconstructions of urban indoor and outdoor scenes designed to evaluate visual locomotion controllers in closed‑loop, continuous testing setups. It aims to bridge the gap between training and real‑world evaluation by providing reproducible environments and policy checkpoints trained across multiple scenes, while highlighting the risks of relying solely on 3D Gaussian‑generated data. The authors offer code and data on their project page for easy integration into robotic evaluation pipelines.
By Ziyu Chen, Henghui Bao, Haoran Chang, Alan Yu, Ran Choi, Kai McClennen, Gio Huh, Kevin Yang, Ri-Zhao Qiu, Yajvan Ravan, John J. Leonard, Xiaolong Wang, Phillip Isola, Ge Yang, Yue Wang
The paper proposes a Friedkin‑Johnsen based framework to identify influential users in online social networks and assess how they shape community opinion. By manipulating initial opinions in experiments, the authors show that top influencers can significantly shift overall community sentiment, and their influence extends beyond direct neighbors to second‑degree contacts. The framework is validated on a tweet dataset from the U.S. presidential election, illustrating the power of digital influencers to alter public opinion.
By Omran Berjawi, Rida Khatoun, Giuseppe Fenza
K-12 robotics and AI education remains difficult to scale, especially in rural regions lacking sustained technical mentorship. Programs like FIRST provide competition pathways and instructional opport...
Interactive control for video generation is moving from coarse prompts toward fine-grained, physically meaningful manipulation of dynamic scenes. Yet existing controllable methods either require the f...
Text-conditioned latent diffusion models perform strongly in video generation and are promising backbones for robotic applications. However, existing approaches rely on pixel-level or VAE-based latent...
In partially observable settings, agents must act without full knowledge of the world state and rely on uncertain state-estimation pipelines. Obtaining grounded and verifiable symbolic plans under suc...
Generative video models can serve as a promising backbone for robot navigation by predicting future observations as video plans. Recent approaches often condition video planning on short-horizon guida...
Learning humanoid-object interaction requires coordinating whole-body balance, locomotion, and dexterous hand contact to control both robot and object motion. Human demonstrations provide examples of...
Zero-shot waypoint navigation requires vision-language models to select, from the current first-person observation, a sequence of spatial actions that is feasible for the agent and reaches the goal, p...
arXiv:2609.13499v1 Announce Type: cross
Abstract: Private Evolution (PE) generates high-fidelity synthetic data in federated settings without exposing users' raw data. It aggregates clipped user vote...
By Sai Aparna Aketi, Enayat Ullah, Shripad Gade
arXiv:2512.15971v2 Announce Type: replace
Abstract: Multispectral object detection is critical for safety-sensitive applications such as autonomous driving and surveillance, where robust perception u...
By Manuel Nkegoum, Minh-Tan Pham, \'Elisa Fromont, Bruno Avignon, S\'ebastien Lef\`evre
arXiv:2609.13152v1 Announce Type: new
Abstract: Large language models (LLMs) perform strongly on static science benchmarks, yet their ability to reason about the physical world through active experim...
By Joseph Chan, Utkarsh Jha, Xiyin Yang, Abhinav Jarajapu, Anik Sahai, Eddie Hu, Robin Jeshua Deepak, Stefano Saravalle, Aditya Shah
arXiv:2609.15277v1 Announce Type: cross
Abstract: Entrepreneurial cognition is a foundation of entrepreneurship research. Yet the growing involvement of large language models (LLMs) in entrepreneuria...
By Christian Fisch, Angela Altmeier, Martin Obschonka, Michal Kosinski, Pin Ni
arXiv:2508.13488v2 Announce Type: replace-cross
Abstract: Loop closure detection is important for simultaneous localization and mapping (SLAM), which associates current observations with historical k...
By Jingwen Yu, Jiayi Yang, Jianhao Jiao, Anjun Hu, Zhonghang Liu, Jiankun Wang, Ping Tan, Hong Zhang
arXiv:2609.13545v1 Announce Type: cross
Abstract: Learning-based adaptive control of robotic manipulators with non-observable friction memory has been addressed by attention- based meta-controllers w...
By Giansalvo Cirrincione, Adriano Fagiolini