arXiv:2610.07982v1 Announce Type: new
Abstract: Monocular metric depth estimation and 3D visual grounding represent the two complementary cornerstones of monocular 3D spatial understanding (M3Sun), f...
By Jinsong Zhang, Kejun Wu, Ming Zhu, Renjie Qiao, Chengtao Cai, Zhengguo Li
arXiv:2610.08068v1 Announce Type: new
Abstract: Realistic physical interaction is a cornerstone of embodied intelligence, yet collecting paired visual--tactile data remains costly. Visual-to-tactile...
By Guo Tang, Yongtao Wang
arXiv:2610.08770v1 Announce Type: new
Abstract: Patch-based learning improves hyperspectral image (HSI) classification by exploiting local spectral-spatial information, but random train-test sampling...
By Mohammed Q. Alkhatib
arXiv:2607.09985v2 Announce Type: replace
Abstract: Object pose estimation is a fundamental problem in 3D vision. Although recent state-of-the-art approaches achieve strong performance, generalizatio...
By Yang You, Yi Du, Cole Harrison, Leonidas Guibas
The paper investigates how rectifying supermarket product images using homography estimation and the Hough transform can improve deep learning-based object detection. It evaluates the impact of angle variation and object density on detection accuracy, highlighting both benefits and limitations of image rectification. The authors advocate for a new dataset to further study these effects.
By Mayank Sah, Jimson Mathew
Con-DSO introduces a consistency-aware RGB‑D direct sparse odometry framework that learns pixel‑level photometric and geometric uncertainty from adjacent RGB‑D frame pairs. The network predicts uncertainties that are converted into pairwise quality scores, guiding support‑pixel selection and forming a host‑side quality prior for keyframe tracking. Experiments on five public benchmarks show that this approach reduces absolute trajectory error by over 20% on ICL‑NUIM and by 50–80% on other datasets, improving robustness in challenging environments.
By Haolan Zhang, Thanh Nguyen Canh, Chenghao Li, Ziyan Gao, Xiongwen Jiang, Nak Young Chong
The paper presents a Vision Transformer-based model for subseasonal soil‑moisture forecasting over Europe, showing that forecast skill depends heavily on how the prediction problem is formulated. By using residual learning and forecasting root‑zone soil moisture in physical units, the model outperforms persistence and existing deep‑learning and ECMWF baselines, providing well‑calibrated probabilistic predictions. However, predicting flash drought onset—defined by rapid multi‑pentad intensification—remains a challenge shared by all current subseasonal‑to‑seasonal systems.
By Noelia Otero, Atahan \"Ozer, Miguel-\'Angel Fern\'andez-Torres, Jackie Ma
WiSPER is a two‑stage framework for multi‑person 3D pose estimation using WiFi channel state information (CSI). The first stage, Pose‑Aware Masked Embedding Learning (PAMEL), couples masked latent prediction with pose‑set supervision to guide the encoder toward joint localization from partial observations. The second stage, Residual Flow refinement with Transformer (ReFT), generates pose candidates for a variable number of people and refines each candidate through a conditional flow guided by coarse coordinates and decoder features. Trained with paired CSI and pose annotations, WiSPER achieves a mean per‑joint position error of 63.72 mm on the PiW3D dataset, a 40.0 % improvement over WiFi‑JEPA and significant reductions for two‑ and three‑person scenarios.
By Gabriel Lee Jun Rong, Shanhong Liu, Pai Chet Ng, Konstantinos N. Plataniotis, Jamal Seyedmohammadi, S. Mohammad Sheikholeslami
The paper introduces a new benchmark dataset of over 6,000 high‑resolution images for detecting and segmenting Western bluebirds in natural, cluttered scenes. Experiments show that supervised detectors such as Faster R‑CNN and RT‑DETR outperform open‑vocabulary models in zero‑shot settings, though fine‑tuned YOLO‑World can compete. Segmentation results favor Mask R‑CNN for mask quality, while YOLOv8‑Seg offers the best precision and speed. The study also identifies multiple factors—scale, brightness, contrast, clutter, blur, crowding, and session variation—as contributors to detection failures.
By Estela Monserrat Arriaga Santana, Julian Rosas Scull, Ibeth P. Alarc\'on, Bibiana Montoya, Aylin Sosa Mej\'ia, Hugo Jair Escalante
The paper introduces FindIt, the first comprehensive benchmark for evaluating the promptable localization abilities of generalist multimodal large language models (MLLMs). It covers four core task categories—object detection, referring expression detection, instance-level detection, and video-based detection—and provides a unified framework that standardizes inputs, enforces parsable bounding box outputs, and defines transparent evaluation protocols. Using this benchmark, the authors assess a range of open-source and proprietary MLLMs, revealing both their strengths and limitations, particularly their sensitivity to formatting constraints and difficulty generalizing to minor variations.
By Eshika Khandelwal, Jingjing Pan, Mingfang Zhang, Quan Kong, Lorenzo Garattoni, Hilde Kuehne
The paper introduces Savvy, a zero‑shot, semi‑online, class‑agnostic system that persistently discovers objects and maintains their identities in long videos, and OGA, an evaluation suite that rewards coherent part‑level predictions even when their granularity differs from reference annotations. Savvy combines modular mask discovery, deferred admission based on accumulated evidence, and track consolidation to keep an evolving object set, outperforming DEVA+SAM and EntitySAM on ScanNet and HM3D datasets in metrics such as VPQ_inf, STQ, and AQ. OGA further distinguishes coherent part‑level support from temporal identity failures, revealing that conventional one‑to‑one VPQ metrics are sensitive to annotation granularity and that temporal failures can be detected even when frame‑level masks remain unchanged.
By Qing Su, Kaiyang Li, Yuan Zhuang, Fei Miao, Shihao Ji
The study introduces a multitask conditional generative adversarial network (MT‑cGAN) that simultaneously synthesizes high‑resolution DESS‑like images and segments knee cartilage and menisci directly from quantitative MRI echo images. Evaluated on 508 knee MRIs from 361 subjects, MT‑cGAN achieved a mean Dice score of 0.84 for segmentation and the lowest coefficient of variation for T1ρ (1.84%) and T2 (1.81%) quantification, outperforming existing conditional GAN approaches. By eliminating the need for separate high‑resolution morphological scans, the method shortens scan times and supports clinical adoption of quantitative MRI.
By Ahmed Tahseen Minhaz, Richard Lartey, Zhiyuan Zhang, Jeehun Kim, Kunio Nakamura, Mingrui Yang, Jiasen Zhang, Weihong Guo, Naveen Subhas, Carl S. Winalski, Xiaojuan Li
The paper reviews uncertainty quantification (UQ) methods for graph neural networks used in connectome-based diagnostic classification and presents a case study on a temporal Graph Attention Network applied to dynamic functional connectivity data for Cocaine Use Disorder. It highlights that deterministic GNNs can produce overconfident predictions, as shown by a Monte Carlo dropout audit revealing high confidence on misclassified subjects. The study demonstrates the need for rigorous UQ, calibration, and selective prediction to ensure reliable graph-based biomarkers in clinical neuroscience.
By Mansooreh Pakravan
arXiv:2610.06942v1 Announce Type: new
Abstract: Deep learning models, particularly recurrent neural networks and their variants, such as long short-term memory, have significantly advanced time serie...
By Nilushika Udayangania, Kishor Nandakishora, Marimuthu Palaniswami
arXiv:2610.07002v1 Announce Type: new
Abstract: Diffusion models learn semantic representations while generating images. In the Decoupled Diffusion Transformer (DDT), a condition encoder provides fea...
By Yiping Ji, James Martens, Simon Lucey
arXiv:2610.08075v1 Announce Type: new
Abstract: Conditional neural fields represent signals continuously, but their effectiveness depends on how the conditional latent representations are inferred fr...
By Rudolf L. M. van Herten, Soufiane Ben Haddou, Rachit Saluja, Johannes C. Paetzold
arXiv:2610.08271v1 Announce Type: new
Abstract: Classical reinforcement learning (RL) assumes that a reward is observed for every visited state-action pair. However, in real-world applications such a...
By Fengxu Liu, Siwei Wang, Gal Dalal, Shie Mannor, Yihan Du
arXiv:2610.07585v1 Announce Type: cross
Abstract: We propose a scalable roto-reflection-group-equivariant vision transformer based on windowed group-convolutional self-attention and a hierarchical fe...
By Sheir A. Zaheer, Jihwan Moon, Chan Y. Park
arXiv:2506.17874v3 Announce Type: replace-cross
Abstract: In many real-world applications, ensuring the robustness and stability of deep neural networks (DNNs) is crucial, particularly for image clas...
By Jiaming Hu, Yeping Jin, Debarghya Mukherjee, Ioannis Ch. Paschalidis
arXiv:2603.27101v2 Announce Type: replace-cross
Abstract: Large-scale maps of field boundaries are essential for agricultural monitoring tasks. Existing deep learning approaches for satellite-based f...
By Gedeon Muhawenayo, Caleb Robinson, Subash Khanal, Zhanpei Fang, Isaac Corley, Alexander Wollam, Tianyi Gao, Leonard Strnad, Ryan Avery, Lyndon Estes, Ana M. T\'arano, Nathan Jacobs, Hannah Kerner