arXiv:2610.00680v1 Announce Type: cross
Abstract: Curvature is often treated as an intrinsic property of a representation, although its empirical effect also depends on coordinate scale, learned logi...
By Athanasios Angelakis, Marta Gomez-Barrero
arXiv:2603.26945v2 Announce Type: replace
Abstract: Appearance-based gaze estimation (AGE) has achieved remarkable performance in constrained settings, yet we reveal a significant generalization gap...
By Zhenhao Li, Zheng Liu, Seunghyun Lee, Amin Fadaeinejad, Yuanhao Yu
arXiv:2607.09086v2 Announce Type: replace
Abstract: We present Subtoken Vision Transformer (SubViT), a selective image tokenization method for fine-grained visual recognition. Standard Vision Transfo...
By Jie Zhu, Ivy Zhang, Minchul Kim, Xiaoming Liu
MIRTO is an evaluation protocol for unsupervised anomaly segmentation in brain MRI that explicitly documents key methodological choices—such as registration alignment, threshold setting, and false‑positive budgeting—and measures their impact. It applies a registration check, uses validation data for thresholding, reports realized false‑positive volumes, and repeats each comparison across 15,552 evaluation pipelines with bootstrap intervals. In a study on four UAD methods and 312 BraTS 2020 subjects, MIRTO revealed that an axis‑order mismatch dramatically lowered a diffusion model’s voxel AUROC, and that many performance differences were driven by lesion definition and threshold transfer rather than model quality.
By Negin Kafee Hernashki, Soumick Chatterjee
Lang3DSeg introduces a point‑transformer backbone for open‑vocabulary, annotation‑free 3D LiDAR segmentation, trained from scratch without geometric pre‑training. It tackles noise from 2D‑to‑3D label projections by applying a class‑priority rule and truncating projected instances at depth gaps, thereby correcting depth‑ambiguity errors. The method achieves state‑of‑the‑art results on nuScenes (52.8 % mIoU) and SemanticKITTI (41.4 % mIoU) while operating in real‑time on a single LiDAR sweep.
By Cigdem Kokenoz, Amir Salarpour, Alkim Domeke, Christopher Salas, Pedram MohajerAnsari, Long Cheng, Mert D. Pes\'e, Bing Li
Dyna‑DINO introduces a curriculum for Vision Transformer (ViT) knowledge distillation that uses the teacher’s intermediate feature maps as progressively harder targets, enabling a student to build foundational representations before tackling higher‑level abstractions. The approach accelerates convergence and improves performance across multiple tasks: on ImageNet‑100 the distilled ViT‑S reaches 90.1% accuracy (+12.24% over baseline), while on ImageNet‑1K it yields +3.9% and +6.09% gains on Oxford and Paris retrieval, +1.93% on semantic segmentation, and notable classification improvements. Additionally, the curriculum reduces training FLOPs by 25.1% and training time by 21% on ImageNet‑100 through early‑stopping of teacher inference.
By Jiaqi Zhang, Ashton Lee, Anthony Wong, John Zou, Sami BuGhanem, Randall Balestriero
LENS‑GRF is a permutation‑invariant lesion evidence network that uses a Set‑Transformer and gated residual fusion to combine global facial context with localized lesion patches for four‑class acne severity grading. The framework integrates adaptive facial skin segmentation, a Vision Transformer prior, and a lesion set transformer that encodes spatial geometry, with a gating mechanism that modulates local residual contributions. In experiments on ACNE04 and PLSBRACNE01, the fully automated model achieved 80.82% accuracy, while using ground‑truth lesion annotations raised accuracy to 95.89% and a Quadratic Weighted Kappa of 0.9753; zero‑shot evaluation on the full cohort yielded 35.00% accuracy versus 42.50% for a global baseline, and oracle analyses on a 148‑subject cohort showed improved accuracy and QWK up to 47.97% and 0.5799.
By Muhammad Muhtasim Shahriar, M. F. Mridha
The paper studies where to place task‑specific adapters in a vision transformer to balance storage growth and accuracy. Training all contiguous four‑block placements shows an inverted‑U accuracy curve, peaking at intermediate depths, while simple weight or activation metrics favor the deepest blocks. A neuroscience‑inspired method, LS‑B, uses frozen fMRI readouts of human visual areas to select blocks whose responses vary most across tasks, yielding backbone‑specific allocations that match or exceed the best placements found by search and use only 60% of the adapter storage while staying within 1.5 percentage points of full accuracy.
By Yuan Huang, Zihan Chen, Runbin Zhang, Hongwei Ding, Changzeng Fu, Shiqi Zhao
arXiv:2610.00279v1 Announce Type: new
Abstract: The segmentation of anatomical structures in medical images and particularly in MRI scans, is essential for clinical diagnosis and monitoring disease p...
By Eirini Cholopoulou, Dimitrios E. Diamantis, Dimitris K. Iakovidis
arXiv:2610.00693v1 Announce Type: new
Abstract: Federated learning (FL) has recently attracted increasing attention in remote sensing (RS) since it enables collaborative model training across decentr...
By Bar{\i}\c{s} B\"uy\"ukta\c{s}, Beg\"um Demir
arXiv:2610.01134v1 Announce Type: new
Abstract: An optical character recognition (OCR) can scan a paper and extract text using technology, making people's jobs easier. While various OCR systems are a...
By Faias Satter, Sk. Md. Masudul Ahsan
arXiv:2604.12102v3 Announce Type: replace
Abstract: We describe compute-grounded reasoning (CGR), a design pattern in which code computes selected sub-problems from explicit intermediate representati...
By Arun Sharma
arXiv:2605.12491v2 Announce Type: replace
Abstract: Vision Transformers (ViTs) learn rich visual-semantic representations through all-to-all self-attention among patch tokens. However, this design im...
By Alan Z. Song, Yinjie Chen, Mu Nan, Deva Ramanan, Michael J. Tarr, Andrew F. Luo
Urban environments are shaped by design choices with long-term implications for health, safety, and quality of life, yet evaluating proposed interventions remains costly, time-consuming, and often imp...
DiDA introduces a lightweight video object segmentation framework that leverages Distillation Learning of Deformable Attention. The method uses deformable attention to adapt key and value positions across frames, enabling object representations that are responsive to spatial and temporal changes. Experiments on DAVIS and YouTube‑VOS benchmarks show state‑of‑the‑art performance and efficient memory usage.
By Quang-Trung Truong, Duc Thanh Nguyen, Binh-Son Hua, Sai-Kit Yeung
OSWorld-Science is a benchmark and evaluation environment for computer-using agents that use visual language models (VLMs) to perform scientific software tasks. It includes 12 VLMs and 146 high-quality tasks across domains such as molecular drawing, pathology image analysis, statistical computing, and physical simulation, with artifact-based evaluation and a harness that logs interactions and supports model comparison. The benchmark was developed through expert proposals and iterative human–AI co‑design, and results show that current VLMs still struggle with key scientific questions, offering insights into factors like language, reasoning, and context length.
By Dingyuan Dai, Heli Qi, Lei Liu, Yinxi Li, Baiding Chen, Zijun Dou, Qingcheng Zeng, Qi Kang, Oliver Sun, Eric Wang, Bo Zhou, Haixin Wang, Yufan Du, Shi Bo, Ruihan Lin, Mengqi Yuan, Dunjie Lu, Steven Dillmann, Yiming Shi, Tina Su, Amy Xin, Minghao Liu, Xi Wang, Xu Huang, Ge Zhang, Pengyu Nie, Zhen Yang, Jie Tang, Juanzi Li, Weihao Xuan, Tianyu Liu
The paper introduces a new hyperspectral image dataset for benchmarking salient object detection, comprising 60 hyperspectral images, their ground‑truth binary masks, and corresponding sRGB renderings. The dataset was curated to include diverse object sizes, counts, contrasts, and positions, addressing the lack of dedicated hyperspectral data for this task. The authors also evaluate existing hyperspectral saliency models using the AUC metric and provide the dataset on GitHub and Hugging Face.
By Nevrez Imamoglu, Yu Oishi, Xiaoqiang Zhang, Guanqun Ding, Yuming Fang, Toru Kouyama, Ryosuke Nakamura
The paper introduces Poincar3, a self‑supervised method that learns multi‑view representations through self‑distillation rather than RGB reconstruction. By combining masked patch and image‑level distillation with a teacher that sees additional views, it trains from scratch without explicit 3D supervision. Poincar3 surpasses prior single‑ and multi‑view self‑supervised methods on tasks such as correspondence estimation, camera pose estimation, and 3D reconstruction, and its features encode camera motion more accurately thanks to a lightweight Poincaré adapter.
By David Nordstr\"om, Thibaut Loiseau, Vincent Lepetit, Michael Felsberg, Guillaume Bourmaud, Fredrik Kahl
arXiv:2609.39265v1 Announce Type: new
Abstract: The Segment Anything Model (SAM) achieves remarkable performance in visual segmentation. The latest SAM3 extends promptable segmentation to concept-lev...
By Ziqi Zhou, Yifan Hu, Yufei Song, Haowen Jiang, Xianlong Wang, Shengshan Hu, Dezhong Yao, Leo Yu Zhang
arXiv:2609.38755v1 Announce Type: new
Abstract: A wide range of approaches have been developed for camera pose estimation, including correspondence-based methods, end-to-end pose regression, and rece...
By Zhining Gu, Shangjie Du, Weimin Qiu, Carl Olsson, Ping Liu, Meng Tang