DualStabSleepNet (DSSNet) is a dual-domain diffusion stabilization network designed to improve the robustness of automatic sleep staging across heterogeneous recording conditions. It employs a continuous-scale diffusion-based module to suppress noise while preserving physiological signals, then transforms stabilized signals into time-frequency representations for a Vision Transformer backbone. A teacher‑student guided diffusion feature stabilization further reduces feature drift, achieving state‑of‑the‑art accuracy on four public PSG datasets and demonstrating strong performance under cross‑dataset distribution shifts.
By Chongjian Wang, Chen Liu, Junjie Gao, Xiaofang Zhong, Shiyuan Han, Tong Zhang
LiAM‑SAM is a lifecycle‑aware memory framework designed to improve segmentation‑based multi‑object tracking (MOT) with the SAM2 foundation video model. It addresses three common failure modes—faulty track initiation, memory drift during close interactions, and unreliable re‑identification after occlusion—by introducing contrastive track initiation, motion‑ and geometry‑grounded memory correction, and adaptive context memory. The system achieves state‑of‑the‑art HOTA and IDF1 scores, with ablations showing significant gains in association metrics and a 96% reduction in identity switches.
By Gr\'egoire Francisco, Alessandro D'Amico, Samuele Costantini, Gianpiero Francesca, Lorenzo Garattoni
The paper introduces a modular perception framework that uses vision‑language models (VLMs) to annotate object‑level regions from a single RGB‑D observation, then grounds these annotations with depth data to build an object‑centric representation. Experiments on 151 tabletop scenes demonstrate that this decomposition maintains strong semantic performance while significantly improving localization and depth estimation compared to direct VLM inference. The resulting representation is integrated into a task‑planning system for robotic manipulation.
By Enrico Saccon, Tommaso Faraci, I\~{n}igo De La Ossa Zarzuelo, Luigi Palopoli, Marco Roveri, Matteo Saveriano
OD3 introduces an optimization‑free dataset distillation framework tailored for object detection. The method first iteratively places object instances in synthesized images, then screens candidates with a pre‑trained observer model to discard low‑confidence objects. Applied to MS COCO and PASCAL VOC, OD3 achieves compression ratios from 0.25% to 5% and surpasses previous detection‑focused distillation methods by over 14% on COCO mAP50 at a 1.0% compression ratio.
By Salwa K. Al Khatib, Ahmed ElHagry, Shitong Shao, Zhiqiang Shen
NV-Reason-CT is a generative vision‑language model designed for chest and abdominal CT analysis that preserves native 3D visual encoding and incorporates radiologist‑guided reasoning. The system couples a 3D vision transformer with a language model, feeding all visual tokens and their 3D coordinates directly into language decoding to maintain volumetric spatial information. Trained on a curated corpus of about 550,000 multimodal instruction examples, the model supports abnormality classification, report generation, and interactive reasoning, achieving strong performance on CT benchmarks and reducing expert interpretation time by 50%.
By Andriy Myronenko, Dong Yang, Yucheng Tang, Baris Turkbey, Benjamin Simon, Stephanie Harmon, Rikhil Makwana, Mariam Aboian, Sena Azamat, Ibrahim Ethem Hamamci, Sezgin Er, Bjoern Menze, Marc Edgar, Yufan He, Pengfei Guo, Daguang Xu
The study investigates whether pathology foundation models (PFMs) carry center-related biases into whole-slide image (WSI) classification. By training models with increasing class-center correlations and evaluating six PFMs across four datasets and two MIL aggregators, the authors introduce the Area Under the Cramér's V Curve (AUCC) to measure both accuracy and degradation due to spurious correlations. Results reveal that center information propagates to WSI predictions, with robustness varying by PFM and MIL strategy, and that ComBat harmonization does not consistently improve robustness.
By Il\'an Carretero, Pablo Meseguer, Roc\'io del Amor, Valery Naranjo
The paper introduces a method for efficiently exploring the Rashomon set of Concept Bottleneck Models (CBMs) by using a parallel parameter‑efficient adaptation module, checkpointing, and a concept diversity objective. This approach generates multiple equally accurate CBMs from a single training process, achieving greater diversity than baseline methods while consuming less memory. The resulting diverse models enable trustworthy selection, reduce inter‑class confusion, and support reliable abstention in decision‑making.
By Shihan Feng, Cheng Zhang, Michael Xi, Ethan Hsu, Lesia Semenova, Chudi Zhong
MessyKitchens introduces a new dataset of cluttered real-world kitchen scenes with detailed 3D object shapes, poses, and accurate contact information. The authors extend the SAM 3D single-object reconstruction method with a Multi-Object Decoder (MOD) to jointly reconstruct entire scenes, achieving better registration accuracy and reduced inter-object penetration compared to prior work. The dataset, benchmark, code, and pretrained models will be publicly released on the project website.
By Junaid Ahmed Ansari, Ran Ding, Fabio Pizzati, Ivan Laptev
arXiv:2609.28239v1 Announce Type: new
Abstract: With the development of applications like autonomous driving, object detection has gained significant attention, while also highlighting critical vulne...
By Li Zeng, Mingcheng Duan, Longfei Fan, Hangtao Zhang, Xianlong Wang, Yanchun Li, Xia Wen, Leo Yu Zhang
arXiv:2609.28283v1 Announce Type: new
Abstract: Several foundation models dedicated to hyperspectral images have recently been made available. These models are trained on large unlabeled datasets and...
By Edgard Dabier, Christophe Kervazo, Pietro Gori, Florence Tupin
arXiv:2609.28327v1 Announce Type: new
Abstract: We present LightMIS, a scalable family of ultra-lightweight convolutional networks for 2D binary medical image segmentation without a learned stage-wis...
By Andrei Arhire, Mihaela-Elena Breab\u{a}n, Radu Timofte
arXiv:2409.13568v3 Announce Type: replace
Abstract: Accurate delineation of agricultural field boundaries is essential for effective crop monitoring and resource management. However, competing method...
By Foivos I. Diakogiannis, Zheng-Shu Zhou, Jeff Wang, Gonzalo Mata, Dave Henry, Roger Lawes, Amy Parker, Peter Caccetta, Suzanne Furby, Rodrigo Ibata, Ondrej Hlinka, Jonathan Richetti, Kathryn Batchelor, Chris Herrmann, Andrew Toovey, John Taylor
arXiv:2609.27988v1 Announce Type: cross
Abstract: Methods operating on Vision Transformer (ViT) feature spaces typically rely on Euclidean distance or cosine similarity. This assumes that every direc...
By Andrew Bond, Ege Erdem \"Ozl\"u, Tuna \c{C}imen, Ilkin Umut Melanlioglu, Tolga Birdal, Erkut Erdem, Aykut Erdem
arXiv:2609.26923v1 Announce Type: cross
Abstract: Cricket is one of the most celebrated sports world-wide, and technological advancement has become deeply embedded in how the modern game is analyzed...
By Sourav Shome, M. D. Ashiquzzaman Rahad, Rameswar Debnath
arXiv:2609.28049v1 Announce Type: cross
Abstract: Video understanding is usually benchmarked on curated, single-actor, or professionally filmed clips, and a strong score there is routinely read as ev...
By Sai Varun Kodathala, Prashanth Pollishetty, Jaylen Cargill
arXiv:2609.28358v1 Announce Type: cross
Abstract: Microscaling quantization techniques are increasingly used to represent neural network parameters with 8 bits or fewer while preserving near-full pre...
By Romain Facq, Sami Ben Ali, Olivier Sentieys
arXiv:2609.27227v1 Announce Type: new
Abstract: Objective assessment of robotic surgery uses instrument kinematics, which must be reconstructed when only video is available. We introduce a kinematic...
By Mehmet Kerem Turkcan, Soham Samal, Zoran Kostic
arXiv:2609.28222v1 Announce Type: new
Abstract: Unified 3D vision-language systems must combine complementary geometry, scale, and appearance cues while supporting tasks from instance segmentation to...
By Xueqi Qiu, Xingyu Miao, Jingjing Deng, Haoran Duan, Yang Long, Ling Shao
arXiv:2609.27094v1 Announce Type: cross
Abstract: Automatic tagging is a core task in Music Information Retrieval (MIR), yet most tagging systems exploit only audio. Live music performance is inheren...
By Alexandros Alexiou, Charilaos Papaioannou, Alexandros Potamianos
arXiv:2604.15221v3 Announce Type: replace-cross
Abstract: Safe human-robot collaboration (HRC) requires accurate human pose estimation and motion prediction to prevent critical collisions. Existing c...
By Jakob Thumm, Marian Frei, Tianle Ni, Matthias Althoff, Marco Pavone