arXiv:2608.30821v1 Announce Type: cross
Abstract: Composable scene modeling aims to recover a real indoor scene as complete, editable object assets arranged as observed, giving robot simulation and e...
By Minghan Qin, Yuang Wang, Xiuyu Yang, Yushi Long, Yujian Zhang, Ruihuan Wang, Kai Ye, Yangang Zhang, Hang Li
The paper reports on training autonomous driving policies via self‑play, extending previous work by using Transformers and a real‑city high‑definition map. On CARLA and Waymo benchmarks, the resulting policies underperform compared to Gigaflow, with identified failure modes such as reward hacking at traffic lights and lack of incentive to stop at stop signs. The authors also analyze which traffic rules emerge from self‑play and confirm that reward conditioning produces diverse driving behaviors.
By Laur Sisask, Ardi Tampuu, Tambet Matiisen
arXiv:2608.29426v1 Announce Type: cross
Abstract: Reliable semantic representations derived from city-scale 3D models are increasingly important for urban analysis, infrastructure monitoring, autonom...
By Alexander Rusnak, Sophia Kovalenko, Jingru Wang, Ismail Moudden, Xiru Wang, Fr\'ed\'eric Kaplan
arXiv:2608.29537v1 Announce Type: cross
Abstract: Frozen vision-language-action (VLA) policies offer broad manipulation skills but execute open-loop action chunks without tracking task progress, so t...
By Hongbo Gao, Zeyu Ni, Xin Wen, Siyu Xu, Ruifeng Li
arXiv:2608.29967v1 Announce Type: cross
Abstract: Vision-Language-Action (VLA) models demonstrate strong semantic understanding yet exhibit systematic failures during deployment. The conditions under...
By Owen Kwon, Pablo Ortega-Kral, Arthur Bucker, Jean Oh
arXiv:2608.30512v1 Announce Type: cross
Abstract: Large-scale industrial robot fleets share constrained physical infrastructure, making vehicle travel times dependent on safety separation, intersecti...
By Cheng Gu, Qiusheng Zhao, Anbang Liu, Shaochong Lin, Max Z. J. Shen
arXiv:2608.31167v1 Announce Type: cross
Abstract: Bridging model-based control and learned policies in long-horizon manipulation has harbored a silent disagreement: control executes specified objecti...
By Weiqi Wang, Zhi Li, Yudong Lei, David Martinez, Xiaofeng Gao, Yuxin Jiang, Chenfanfu Jiang, Yingnian Wu, Demetri Terzopoulos, Ran Gong
arXiv:2503.13415v2 Announce Type: replace-cross
Abstract: With the rapid development of artificial intelligence, intelligent decision-making techniques have gradually surpassed human levels in variou...
By Weiqiang Jin, Hongyang Du, Shixiang Tang, Biao Zhao, Guang Yang
arXiv:2608.28664v1 Announce Type: cross
Abstract: Despite the success of foundation models in language and vision, their expansion into embodied AI is bottlenecked by a lack of generalized touch sens...
By Jianquan Wang, Haiwei Dong, Abdulmotaleb El Saddik
The paper introduces a fully distributed continuous‑time algorithm for solving Generalized Nash Equilibrium Problems (GNEPs) with shared linear equality constraints. Unlike existing methods that require exchanging Lagrange multipliers, this approach converges to any GNE without multiplier communication, thereby reducing communication overhead and enhancing privacy. Discrete‑time variants are also presented and the method is demonstrated on a multi‑robot placement task.
By Shao-An Yin, Mingyi Hong, Nicola Elia
arXiv:2508.13073v3 Announce Type: replace-cross
Abstract: Robotic manipulation, a key frontier in robotics and embodied AI, requires precise motor control and multimodal understanding, yet traditiona...
By Rui Shao, Wei Li, Lingsen Zhang, Renshan Zhang, Zhiyang Liu, Ran Chen, Liqiang Nie
RoboPhys-3D is a 3D‑grounded embodied world model benchmark built on RoboTwin 2.0, featuring 50 manipulation tasks, 5,000 episodes, and 25,000 multi‑view ground‑truth videos. It evaluates video world models by processing both generated and ground‑truth videos through the same 3D reconstruction pipeline, allowing the separation of reconstruction‑induced from generation‑induced errors. The benchmark defines 50 metrics across four sub‑dimensions—pixel fidelity, 3D geometry consistency, state understanding, and task completeness—and introduces the Average Full Score and RoboPhyscore for holistic assessment, with RoboPhyscore showing strong correlation with human judgments.
By Tianyi Wang, Jiazhou Chen, Yiming Xu, Xiangyu Li, Tianyi Zeng, Chih-Hsien Chou, Ning Lu, Liang Peng, Junfeng Jiao, Christian Claudel
The study investigates how multimodal large language models (MLLMs) use prosodic cues in sarcasm detection. By testing Qwen2.5‑Omni and Qwen3‑Omni on Mandarin Chinese and English across five modality conditions, the authors find that adding audio increases false positives without improving true positives. Acoustic error analysis shows that models rely on a stereotypical prosodic pattern—elevated pitch and irregular pausing—that does not align with genuine sarcasm cues, and manipulating these dimensions alone can raise false positive rates up to 60%. The same effect appears in Gemini 3 Flash Preview, indicating the heuristic is not limited to a single architecture.
By Yongjian Chen, Pengfei Wei, Yiqun Sun, Zhu Li, Lawrence B. Hsieh
arXiv:2608.28948v1 Announce Type: cross
Abstract: Over the course of the last decade, neural networks have grown from an academic curiosity to moving the markets of nations. Despite this explosion in...
By David Aram Yunis
arXiv:2608.30714v1 Announce Type: new
Abstract: Verifying image authenticity is increasingly difficult, posing serious risks across journalism, law enforcement, and political domains. Most existing f...
By Siddhi Pravin Lipare, Vishesh Kumar, Akshay Agarwal
arXiv:2608.30400v1 Announce Type: new
Abstract: High-dynamic-range (HDR) images, with their rich tone and detail reproduction, hold significant potential to enhance computer vision systems, particula...
By Gongzhe Li, Linwei Qiu, Peibei Cao, Fengying Xie, Xiangyang Ji, Qilin Sun
arXiv:2608.30344v1 Announce Type: new
Abstract: Most shape reconstruction methods assume measurements defined over planar sensing domains, such as RGB images or depth maps. In this paper, we use a cu...
By Hao Chen, Chenming Wu, Chun Ping Lam, Xiangjia Chen, Guoxin Fang, Charlie C. L. Wang, Yeung Yam, Juncong Lin, Chengkai Dai
The paper introduces OHCAM, an online method for learning action models that include conditional and quantified effects from limited interactions. It maintains a belief over possible models and actively chooses actions that maximize disagreement among hypotheses to reduce uncertainty, while handling noisy observations. Starting with simple hypotheses, OHCAM expands complexity only when necessary, achieving sample‑efficient learning that outperforms baselines on benchmark domains and is validated on a Kinova Gen3 robot.
By Jeffrey Jewett, William Solow, Sandhya Saisubramanian
arXiv:2608.29475v1 Announce Type: new
Abstract: Surface material recognition from incomplete visual observations remains a challenging problem in robotic perception and environmental understanding. T...
By Sindhuja Penchala, Sudip Mittal, Noorbakhsh Amiri Golilarz
arXiv:2608.29937v1 Announce Type: new
Abstract: Latent world-action models avoid rendering future pixels by predicting an action-relevant visual subgoal in feature space. LaWAM established this formu...
By Yafei Zhang, Nan Wu