The paper introduces BIFTA, a Brain‑Inspired Few‑Shot Tactile Adaptation framework that enables a frozen encoder to adapt quickly to an unknown tactile sensor using only a small labeled support set. It preserves pretrained representations via dual‑view statistical memory, builds support‑conditioned spectral graphs to correct sensor‑dependent feature neighborhoods, and employs uncertainty‑gated recurrent propagation to reinforce reliable cross‑query evidence. Benchmarks on three tactile datasets demonstrate that BIFTA dramatically improves adaptation performance, achieving an 87.09% mean Sparsh accuracy on SITR with just 10% labeled data—an increase of 47.22 percentage points over the best prior method.
By Boheng Liu, Ziyu Li, Xia Wu
The paper investigates adding Greek language support to a Cosmos3 vision‑language‑action robot policy using only machine‑rephrased instructions and no architectural changes. It finds that many evaluation metrics can give misleading results, and that multilingual training with Greek yields a modest 6.7‑7.1 point advantage over a control, reaching about 40% of English performance. The study also shows that overfitting to a single translator’s phrasing can be mitigated by training on multiple phrasings, while warm‑starting from a language‑adapted world model or unfreezing the text tower actually harms performance.
By Ayoub Kirouane, Georgios Giaples, Christos Petrocheilos
The paper investigates how reinforcement‑learning policies for legged robots encode gait information by examining the effective rank of the policy Jacobian conditioned on gait phase. It finds that common architectural features such as layer normalization and residual connections allocate more representational capacity to swing than stance, a pattern absent in vanilla MLPs. Leveraging these insights, the authors propose a simple recipe that improves sim‑to‑real transfer, reducing joint jitter on a physical Spot robot by roughly three‑fold.
By Felipe Tommaselli, Thiago H. Segreto, Juliano D. Negri, Ricardo V. Godoy, Marcelo Becker
The paper presents a multi‑modal deep learning model that uses temporal attention to detect internal welding defects such as porosity, lack of penetration, fusion, undercut, and cold lap in fillet joints during real‑time Gas Metal Arc Welding. Trained on images and sound data from an industrial collaborative welding robot, the model achieves an F1 score of 0.99. Explainable AI techniques are applied to interpret the model’s behavior, highlighting key image and sound spectrogram regions and the most effective modality for each defect type, thereby enhancing trust and reliability in AI‑driven welding inspection.
By Mobina Mobaraki, Mahyar Asadi, Klaske Van Heusden, Guy A. Dumont
CogniDir is an adaptive distributional learning framework designed to improve fake news detection against new psychologically grounded malicious comments generated by Large Language Models. It reframes robust detection as a dynamic data mixture optimization problem, using cognitive psychology to formalize adversarial paradigms and an information‑theoretic score to guide adaptive sampling of training data. Experiments on three benchmarks show that CogniDir achieves state‑of‑the‑art robustness, boosting F1 scores by up to 17.9% over existing baselines under heterogeneous AI‑generated attacks.
By Zhao Tong, Chunlin Gong, Yimeng Gu, Haichao Shi, Qiang Liu, Shu Wu, Xingcheng Xu, Xiao-Yu Zhang
ARC‑Bench is a new benchmark that tests whether frozen JEPA‑style latent world models can correctly rank candidate actions by latent distance. The study finds that the assumption of latent rankability fails dramatically in both navigation and manipulation tasks, with the top‑scored actions often being suboptimal. Closed‑loop replanning masks this defect, but reducing replanning frequency reveals the underlying ranking failures.
By Zhengshu Zhang, Zhiyuan Li
arXiv:2609.08084v1 Announce Type: cross
Abstract: Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computati...
By Igor Pavlovic, Thiemo Wandel, Anton Obukhov, Luca Bartolomei, Andrey Davydov, Fabio Tosi, Matteo Poggi, Sabine S\"usstrunk, Dengxin Dai
The paper studies how Bird's‑Eye‑View (BEV) maps predicted by Cross‑View Transformers (CVT) can be used directly as inputs to a Behavior‑Cloning (BC) driving policy in the CARLA simulator. It introduces a six‑channel BEV representation and a Kernel Density Estimation (KDE) weighting scheme to focus learning on underrepresented maneuvers. Closed‑loop tests show that the KDE‑weighted model is the only predicted‑BEV agent to finish an episode without infractions, highlighting that global segmentation scores are poor proxies for driving performance and that prediction quality at critical geometries, especially the route channel, is key to reliable navigation.
By Felipe Carlos dos Santos, Eric Antonelo, Gustavo Claudio Karl Couto
The paper presents a solution for the UCF UrbanTwin LUMPI Track in the Sim-to-Real LiDAR Challenge, where a detector trained solely on synthetic data must perform on real LiDAR frames. The approach tackles the Sim2Real gap through data alignment, diversified sampling, augmentation, and specialized detectors, followed by class-aware fusion and calibration techniques. The final submission achieved a Combined Score of 0.4692, a Detection Score of 0.1797, a Realism Score of 0.9035, and a 3D mAP@0.5 of 0.1258.
By Pu Luo, Cong Xu, Yumei Li, Kexin Zhang, Licheng Jiao, Wenping Ma, Lingling Li
The paper introduces a structure‑aware federated learning framework for segmenting catheters and guidewires in X‑ray fluoroscopy. It presents a new benchmark dataset, CathAction, and a shape‑sensitive loss that improves segmentation accuracy. The approach extends to federated learning, adding projected gradient descent for adversarial optimization, and includes a diffusion‑based synthetic data generator that boosts performance under data scarcity.
By Chayun Kongtongvattana
D3ARC is an asynchronous distributed hierarchical framework designed for time‑critical wildfire detection using multiple robotic agents. It enables cooperative perception, shared situational awareness, and coordinated actions while a remote controller asynchronously directs each robot’s motion. The system incorporates safe navigation, coverage efficiency, and a forward‑looking capability to evaluate candidate strategies before execution, achieving up to 94% mission success and 89.4% detection confidence in realistic simulations.
By Nikolaos Koursioumpas, Lina Magoula, Nancy Alonistioti, Ramin Khalili
The article titled "The convergent laboratory: when AI reasoning, autonomous experiments, high performance and quantum computing reshape chemistry" discusses insights from the TPC26 conference, where leaders from academia, national laboratories, and industry examined how AI, autonomous agents, self-driving labs, high‑performance computing, and quantum computing converge to accelerate materials science discovery. It presents firsthand experiences from researchers at the forefront of these technologies and argues that their simultaneous maturation marks a tipping point for transformative advances and productive disruption in chemical sciences.
By Eliu Huerta, Xiaoyun Wang, Geetika Gupta, Edward H. Sargent, Cameron J. Owen, Victor Fung, Abhishek Mitra, Austin Cheng, Emma Bouchard, Shams Mehdi
The paper presents a deep active inference framework for real‑world robotic navigation that combines a diffusion policy with a multiple‑timescale recurrent state‑space model. The diffusion policy generates diverse candidate actions, while the state‑space model predicts long‑horizon outcomes, allowing the system to select actions that minimize expected free energy. Experiments show higher success rates and fewer collisions, especially in exploration‑heavy scenarios, demonstrating the effectiveness of this unified exploration and goal‑directed approach.
By Riko Yokozawa, Kentaro Fujii, Yuta Nomura, Shingo Murata
The paper introduces a progressive training strategy for embodied vision‑language models aimed at reducing spatio‑temporal hallucinations. It first creates a Chain‑of‑Thought dataset that breaks complex reasoning into detailed spatiotemporal steps, then uses supervised pre‑training on this dataset followed by fine‑tuning with weakly‑labeled data. Experiments show the method improves backbone accuracy and narrows the forward‑backward performance gap from over 70% to 6.53%, indicating stronger dynamic reasoning and fewer temporal biases.
By Xiaoda Yang, Shuai Yang, Can Wang, Jingyang Xue, Menglan Tang, Checheng Yu, Xunzhe Zhou, Sashuai Zhou, Tao Jin, Lixin Yang, Xiangyu Yue, Zhou Zhao
The paper "Inferring the Unspoken: Aligning Embodied Agents with Implicit Preferences" addresses the challenge of natural-language instructions that omit details needed for embodied action. It introduces the Preference-based Planning (PbP) benchmark, comprising 5,000 evaluation groups and 290 preferences across three levels, to systematically evaluate agents’ ability to infer latent user preferences from a few demonstrations. The authors propose the two-stage Inferring the Unspoken (InTU) framework, which first verbalizes inferred preferences from multimodal demonstrations and then generates action plans conditioned on that explicit representation, showing that explicit verbalization improves alignment and robustness compared to direct end-to-end planning.
By Manjie Xu, Xinyi Yang, Wei Liang, Chi Zhang, Yixin Zhu
PlannerForge is a unified LLM‑agent framework that covers the entire scenario‑based testing pipeline for autonomous driving systems, from scenario generation to ADS assessment, and adds ADS enhancement and benchmarking stages. It was evaluated with ten off‑the‑shelf LLMs across all tasks and five prompt conditions, achieving best‑per‑task scores between 0.88 and 1.00 and matching commercial APIs with open‑source models such as Qwen3.6:35B. The end‑to‑end chaining retains 83% of seed queries for commercial backends and 78% for open‑source, outperforming existing tools like Scenario Factory 2.0 and BM25 in natural‑language generation, attribute realization, and physically valid edits.
whyItMatters":"PlannerForge demonstrates that a single LLM‑based system can streamline and improve the fragmented scenario‑based testing workflow for autonomous driving, achieving high performance without domain‑specific fine‑tuning."
By Yuan Gao, Sebastian M\"uller, Mattia Piccinini, Marc Kaufeld, Yuchen Zhang, Finn Rasmus Sch\"afer, Qunying Song, Johannes Betz
RoboCousin is an extensible simulation platform that transforms user-provided observations into reusable assets, scenes, and expert trajectories for bimanual robotic manipulation. It converts object images into simulation-ready models with visual, collision, semantic, and physical metadata, automatically generates grasp candidates, and builds digital cousins that vary objects, backgrounds, layouts, and language instructions while preserving task-relevant affordances. The platform supports both tabletop and room-level scene construction, and the authors release RoboCousin-OBD with over 3,000 annotated objects and 50 backgrounds, generating more than one million expert trajectories across 50 tasks, with simulation and real-robot experiments demonstrating comparable annotation quality and effective sim-to-real transfer.
By Jingxuan Zhu, Jingyi Li, LiangLiang Chen, Zhiyuan Jing, Jidong Zhang, Hongming Li
The paper introduces a multi‑modal late‑fusion perception pipeline for object detection and tracking in autonomous racing. It combines independent detections from cameras, LiDARs, and RADARs to produce timely and robust state estimates of surrounding vehicles. The tracking framework compensates for detection delays and incorporates vehicle dynamics and track layout knowledge, and its effectiveness is confirmed through real‑world experiments in diverse critical scenarios.
By Davide Malvezzi, Michele Pestarino, Vittoria Cavicchioli, Valentina La Gamba, Silvia Severi, Fabio Bagni, Luca Bartoli, Massimiliano Bosi, Francesco Gatti, Micaela Verucchi, Ayoub Raji, Marko Bertogna
FALCON‑S is a modular, high‑fidelity simulator designed for fixed‑wing aerial robots operating near the ground. It models full 6DoF rigid‑body physics, semi‑empirical ground‑effect aerodynamics, actuator dynamics, sensor noise, and environmental disturbances, and supports both CPU and GPU backends via Torch and NVIDIA Warp for large‑scale reinforcement learning and optimal control. The framework offers a unified interface for various controllers, including RL and optical control algorithms, and allows cross‑validation with X‑Plane and JSBSim for engineering integration and visual fidelity.
By Matteo El Hariry, Pedro Lima, Andrej Orsula, Matthieu Geist, Miguel Olivares-Mendez
PV-WM is a history‑only world model that jointly predicts pedestrian root motion, 15‑joint articulation, and vehicle kinematic states in a synchronized heterogeneous state. It uses recurrent updates to generate pedestrian and vehicle motion chunks, reconstructing vehicle boxes from predicted center, heading, and observed extent, and recomputes pedestrian‑vehicle geometry after each transition. Compared to a one‑shot predictor, PV‑WM reduces Root ADE by 12.7% and MPJPE by 14.8%, and across 824 Waymo contexts it lowers Root ADE by 5.2%, MPJPE by 7.6%, P‑V distance error by 11.9%, and oriented‑box closest‑approach error by 5.8%, while using 57.1% fewer parameters, 96.5% fewer FLOPs, and 25.5% lower p95 latency.
By Haozhuang Chi, Jingsong Liang, Ziying Song, Lei Yang, Shihao Li, Haoruo Zhang, Chen Lv