arXiv:2610.09718v1 Announce Type: cross
Abstract: Vision-Language-Action (VLA) models acquire broad manipulation capabilities via large-scale pretraining, yet eliciting them through language requires...
By Masatoshi Tateno, Takehiko Ohkawa, Yueh-Hua Wu, Hanlong Li, Tatsuya Matsushima, Yoichi Sato, Kei Ota
ActiveLang is an autonomous system that builds language‑annotated 3D maps by actively exploring environments. It uses a compact dual‑Gaussian representation to jointly reconstruct geometry, appearance, and open‑vocabulary semantics while adapting language features online. The planner selects informative viewpoints, reducing the number of observations and computational cost, and experiments on Replica and ScanNet++ show significant gains in 2D and 3D open‑vocabulary segmentation compared to existing baselines.
By Liyan Chen, Hairong Yin, Huangying Zhan, Yi Xu, Raymond A. Yeh, Philippos Mordohai
WAPR is a zero‑shot wide‑angle pose refinement model that can correct candidate 6D poses with rotational errors up to 90°, achieving fast inference (≤1 s per frame) and high throughput (≈25 detections per second). It leverages rotational symmetry priors to canonicalize pose targets and introduces the SA6D dataset, which augments 944 GSO scans into ~50 k object instances and ~2 M RGB‑D images. Experiments on seven BOP core datasets demonstrate that WAPR sets new state‑of‑the‑art performance for unseen‑object pose estimation in both fast and unconstrained settings.
By Yulin Wang, Mengting Hu, Hongli Li, Jianghao Zhou, Chen Luo
The paper introduces a lightweight learned optimizer that recombines gradient history by averaging over disjoint time spans, reducing the prediction space to a single scalar coefficient per average shared across parameters. By progressively averaging older gradients, the method keeps memory usage low while maintaining independent contributions from long‑term history. A small 37k‑parameter network trained in under a GPU‑hour generalizes zero‑shot to unseen tasks, improving validation loss on BERT‑Tiny and GPT‑Tiny and boosting test accuracy over Adam on Vision Transformers and graph models with minimal FLOPs overhead.
By Minyoung Choi, Dalta Imam Maulana, Wanyeong Jung
The paper introduces a method to estimate a model’s vulnerability to the Likelihood Ratio Attack (LiRA) without training reference models, using only the target model’s train and test loss distributions. It shows that LiRA’s per‑sample signal can be decomposed into a variance‑ratio term and a residual mean‑shift term, and that different loss‑distribution shapes dictate which reference‑free proxy to use. Two proxies are presented: the LOSS attack TNR for heavy‑tailed losses and the LOSS attack AUC for symmetric losses, both achieving low RMSE in predicting LiRA TPR across multiple architectures and datasets.
By Euodia Dodd, Nata\v{s}a Kr\v{c}o, Igor Shilov, Matthew Wicker, Yves-Alexandre de Montjoye
LiG-DETR introduces a Global-Local Reassembly framework for aerial object detection that captures high‑fidelity local features before compression and integrates them into a unified end‑to‑end DETR decoder. The method uses a shared encoder to extract both global and locally magnified features, reassembles the local features according to their spatial positions, and employs Context‑Preserved Selective Reassembly and Density‑Aware Adaptive Query Allocation to reduce redundant computation. Experiments demonstrate significant improvements on small and medium objects while maintaining strong performance on large objects, with favorable accuracy–efficiency trade‑offs and better cross‑domain generalization.
By Yupeng Zhang, Fangzhuo Gao, Juntao Cheng, Ziyi Zhao, Liang Wan, Ruize Han
arXiv:2604.08015v3 Announce Type: replace-cross
Abstract: Small lesions in brain MRI are hard to segment because they occupy a tiny fraction of the volume and are dominated by background and larger l...
By Minh Sao Khue Luu, Evgeniy N. Pavlovskiy, Bair N. Tuchinov
arXiv:2610.08794v1 Announce Type: new
Abstract: Large language models rely on subword tokenizers whose quality varies across languages, yet no standardized multi-metric framework exists for broad com...
By Ben Gubler
arXiv:2608.16303v2 Announce Type: replace
Abstract: Long-term emotional-support agents require memory mechanisms for personalized understanding across sessions. However, emotional-support dialogue is...
By Chang Liu, Junyi Zhao, Shuyi Zhang, Changsheng Ma, Yongfeng Tao, Minqiang Yang, Bin Hu
arXiv:2610.08936v1 Announce Type: new
Abstract: Robustness of image classification has several benchmarks, but their video counterparts are absent. In video classification temporal dimension introduc...
By Maksim Plinskiy, Aleksandr Gushchin, Sergey Lavrushkin, Dmitriy S. Vatolin, Anastasia Antsiferova
arXiv:2610.09185v1 Announce Type: new
Abstract: Cine cardiovascular magnetic resonance (CMR) captures the cardiac cycle as a four-dimensional (4D) sequence, but standard acquisition requires electroc...
By Shiyi Wang, Ruochen Sun, Xiang Li, Peirong Liu, Fangxu Xing
arXiv:2610.09392v1 Announce Type: new
Abstract: Reliable volumetric segmentation is critical for clinical diagnostics, yet foundation models such as MedSAM remain deterministic and lack calibrated un...
By Shadi Alijani, Fereshteh Aghaee Meibodi, Homayoun Najjaran
arXiv:2610.09397v1 Announce Type: new
Abstract: Cine cardiovascular magnetic resonance (CMR) analysis relies on multi-frame sequences capturing the full cardiac cycle. However, standard multi-frame a...
By Shiyi Wang, Ruochen Sun, Peirong Liu, Xiang Li, Fangxu Xing
arXiv:2610.09439v1 Announce Type: new
Abstract: Ancient script image restoration is a fundamental problem in computer vision, as it directly affects the reliable analysis and interpretation of histor...
By Jaidev Sanjay Khalane, Akbar Ali, V. N. Prabhakar, Shanmuganathan Raman
arXiv:2610.09455v1 Announce Type: new
Abstract: Recently, approaches that leverage human video datasets for robot policy training have become increasingly prevalent. However, most existing hand track...
By Seungjun Moon, Subin Jeon, Sangwoo Kim, Hanbyul Joo, Jinwoo Shin
arXiv:2610.09534v1 Announce Type: new
Abstract: Rotational symmetry is an important prior in 6D pose estimation, improving pose accuracy and supporting symmetry-aware evaluation. However, current sym...
By Mengxin Zhang, Yulin Wang, Chen Luo, Yongzhe Li, Yijun Zhou
arXiv:2610.09849v1 Announce Type: new
Abstract: Selecting synthetic degradations for dense prediction requires an estimate of their training utility, the generalization gain they bring under a finite...
By Chunming He, Kailai Zhou, Jiaming Zuo, Hanqi Liu, Fengyang Xiao, Youwei Pang, Xiaofeng Liu, Weisi Lin, Xiaoqi Zhao
arXiv:2610.09907v1 Announce Type: new
Abstract: Vision-Language Models (VLMs) have shown strong performance across a wide range of downstream vision tasks, thanks to the complementary information con...
By Ankita Das, Ambarish Parthasarathy, Sumohana S. Channappayya, C. Krishna Mohan
arXiv:2610.10012v1 Announce Type: new
Abstract: In the framework of edge-weighted graphs, watersheds have proven to be linked to well-known optimization problems, as Minimum Spanning Tree, which allo...
By Jean Cousty (LIGM), Laurent Najman (KUSTAR, LIGM), Benjamin Perret (LIGM), Deise Santana Maia (CRIStAL)
arXiv:2610.10116v1 Announce Type: new
Abstract: Semantic segmentation networks operate on a fixed set of classes and therefore fail when out-of-distribution (OOD) objects appear during deployment, a...
By Arnold Brosch, Abdelrahman Eldesokey, Michael Felsberg, Kira Maag