arXiv:2609.39083v1 Announce Type: new
Abstract: Super-resolution and quality enhancement of 1.5\,T brain MRI are normally validated with image-fidelity metrics, although their purpose is to improve d...
By Kavitha Viswanathan, Harsh Choudhary, Amit Sethi
arXiv:2609.39184v1 Announce Type: new
Abstract: Fiber orientation and compartmental microstructure are central to the characterization of white matter tissue in diffusion MRI, yet existing methods ei...
By Sebastian Endt, Marcus Wirth, Johannes Reinhold Schlund, Marion Irene Menzel
arXiv:2609.39467v1 Announce Type: new
Abstract: Pedestrian detection plays a crucial role in computer vision with applications in autonomous driving, surveillance, and public safety. However, real-wo...
By ZiAn Wang, MingZhe Liu, Chaoyi Guo, ChangChun Li, Fangming Gu
arXiv:2609.39627v1 Announce Type: new
Abstract: This book presents a code-first introduction to computer vision, spanning classical 2D image processing, classical 3D vision, and deep learning. Organi...
By Stan Birchfield
arXiv:2609.39681v1 Announce Type: new
Abstract: Unsupervised domain adaptation (UDA) reduces the annotation burden in panoptic segmentation by leveraging a cost-effectively labeled source domain (e.g...
By Ivan Martinovi\'c, Josip \v{S}ari\'c, Yuki M. Asano, Sini\v{s}a \v{S}egvi\'c
arXiv:2609.39785v1 Announce Type: new
Abstract: The Segment Anything Model (SAM) relies heavily on massive manual annotations, creating a fundamental bottleneck for model scaling. While unsupervised...
By Weijian Jian, Xiaoyue Zhang, Bin Xiao, Chunyu Xie, Yixiao He, Yutao Liu, Dawei Leng, Yuhui Yin
arXiv:2609.40007v1 Announce Type: cross
Abstract: A vision-language-action (VLA) policy can finish a manipulation task while knocking over objects unrelated to it, so task success alone does not show...
By Yatharth Agarwal, Vijay Raghunathan
arXiv:2512.22819v2 Announce Type: replace
Abstract: Panoramic depth estimation captures the complete 360$^\circ$ scene geometry, being essential for robotics and AR/VR applications. While perspective...
By Hualie Jiang, Ziyang Song, Zhiqiang Lou, Rui Xu, Minglang Tan
arXiv:2603.27519v4 Announce Type: replace
Abstract: Image-based plant phenotyping depends on dense structural understanding of crops, yet pixel-level annotation remains expensive across species, orga...
By Shuai Xiang, James Burridge, Shouyang Liu, Hao Lu, Tokihiro Fukatsu, Yinqiang Zheng, Wei Guo
arXiv:2605.17630v3 Announce Type: replace
Abstract: Frozen segmentation foundation models often fail when the target class appears in a form that is weakly represented during pretraining. To address...
By Abderrahmene Boudiaf, Irfan Hussain, Sajid Javed
arXiv:2609.34330v2 Announce Type: replace
Abstract: Multimodal large language models (MLLMs) have demonstrated impressive performance in multimodal understanding, but processing large numbers of visu...
By Tinghao Wang, Yichen Guo, Qizhe Zhang, Yuan Zhang, Weimin Ouyang, Rui Huang, Jiajun Cao, Sixiang Chen, Hao Jiang, Jixian Wu, Zheng Lu, Bofan Zhu, Renyuan Li, Shanghang Zhang
Cryo-Bench is a new benchmark that evaluates foundation models for cryosphere mapping, comprising six semantic‑segmentation datasets across five cryospheric components (supraglacial debris, glacial lakes, sea ice, calving fronts, and Antarctic ice‑shelf extent). The benchmark includes multispectral, RGB, and SAR observations from under‑represented regions and tests thirteen geo‑foundation models alongside U‑Net and Vision Transformer baselines. Results show that with frozen encoders U‑Net slightly outperforms TerraMind, but the difference is not statistically significant; fine‑tuning with learning‑rate optimization can dramatically improve performance for some models, while in few‑shot scenarios several foundation models retain over 90 % of their full‑label accuracy.
By Saurabh Kaushik, Lalit Maurya, Beth Tellman, Swalpa Kumar Roy, Valerio Marsocci, Gustau Camps-Valls, Jocelyn Chanussot
Mobile robots and vehicles carry synchronized multi-camera rigs, yet many streaming 3D foundation models are designed for monocular input, leaving efficient use of rig geometry a challenge. We present...
This paper demonstrate that whether masking-based token pruning helps or hurts worst-group robustness can be predicted before deployment, without labels or fine-tuning. A systematic study of semantic...
The paper presents an OCR model designed to extract key information from Korean-language invoices. It combines deep learning with image preprocessing techniques and achieves an 87% F1‑score on a diverse set of collected invoices while maintaining negligible processing time.
By Xiem HoangVan, Phu TranQuang, Minh DinhBao, Tien VuHuu
HyperSAM is a promptable foundation model for hyperspectral remote sensing that integrates a data‑centric synthesis pipeline with a spectral adaptation architecture based on Segment Anything Model 3 (SAM3). The model generates full‑spectrum hyperspectral cubes from high‑resolution multispectral imagery using a physics‑informed abundance‑transfer generator, and employs SAM3‑derived pseudo‑masks for object‑centric supervision. With a frozen SAM3 RGB branch, a trainable hyperspectral encoder, ControlNet‑style feature injection, and a mixture‑of‑experts mask refiner, HyperSAM demonstrates strong generalization across diverse hyperspectral tasks such as classification, anomaly detection, change detection, target detection, and airborne oil‑spill mapping.
By Li Pang, Xinqiao Wu, Jing Yao, Pedram Ghamisi, Jun Zhou, Zhengchao Chen, Deyu Meng, Xiangyong Cao
The paper introduces a controlled benchmark for evaluating large language models (LLMs) on key‑value pair extraction from documents with varying levels of OCR noise. It tests 136 configurations across five instruction‑tuned open‑weight LLMs, three datasets, and four text‑quality conditions, using deterministic decoding to generate 17,688 document‑level inferences. The study finds that clean‑text performance does not reliably predict real‑world robustness, model rankings can reverse under noisy conditions, and few‑shot demonstrations do not always improve accuracy, highlighting reliability risks in OCR‑to‑LLM pipelines.
By Zahra Anvari
Merlin Plus is a new, large-scale CT dataset that provides radiologist‑created tumor masks for nine different organs, adding 1,153 per‑voxel masks and longitudinal metadata to the existing Merlin collection. The dataset was built using a report‑based active‑learning framework, where radiology reports flag tumor cases, a segmentation model generates initial masks, and radiologists review and correct them, thereby reducing annotation effort while preserving high quality. The added longitudinal data enables temporal modeling of cancer progression, supporting scalable multi‑organ cancer detection, segmentation, and longitudinal analysis in CT.
By Pedro R. A. S. Bassi, Wenxuan Li, Szymon Plotka, Ruby Honjol, Jakub Przado, Xinze Zhou, Kang Wang, Yang Yang, Malte Jensen, Akshay S. Chaudhari, Curtis P. Langlotz, Alan L. Yuille, Zongwei Zhou
arXiv:2609.36969v1 Announce Type: new
Abstract: 3D Gaussian Splatting (3DGS) is a state-of-the-art technique for 3D scene rendering, offering high efficiency and excellent visual quality. However, be...
By Gyeonggwan Lee, Seunghwan Hong, Junghun Suh
ByteTraX is a lightweight enhancement to the ByteTrack multi‑object tracking architecture that introduces a single unified matching threshold and stricter track initiation criteria to reduce erroneous track reclassification and identity switches. The method yields consistent performance gains across several benchmarks—GMOT‑40, LC‑MOT, SportsMOT, TeamTrack, DAMUNT, and DeepSea‑MOT—while boosting processing speed by over 10%. Quantitatively, ByteTraX achieves more than a 40% drop in identity switches, with mean improvements of 3.6 in HOTA, 5.6 in IDF1, and 6.3 FPS.
By Thomas A. O'Shea-Wheller