The paper introduces Quantum-Inspired Nonlinear Adapters (QINA), compact modules that apply learnable trigonometric feature lifting followed by bounded nonlinear aggregation to pretrained vision models. QINA enables structured oscillatory basis functions with a norm-dependent Lipschitz bound, allowing spectral reshaping of representations without expanding the receptive field or significantly increasing parameters. Experiments on natural and medical imaging tasks show that QINA consistently outperforms identity baselines, fixed Fourier mappings, and parameter-matched generic adapters, demonstrating that geometry- and spectrum-aware adaptation is crucial for effective frozen-backbone transfer learning.
By Mostafa Mehdipour Ghazi
The paper introduces Selective Probability Mass Concentration (sPMC), a training framework that strengthens implicit visual grounding in multimodal large language models by selectively regularizing attention heads most responsive to visual evidence. sPMC treats attention over visual tokens as a spatial probability distribution and encourages mass to concentrate on semantically relevant regions using segmentation-derived priors, while leaving other heads unconstrained. Across six multimodal benchmarks, sPMC yields an average zero‑shot improvement of 3% and gains up to 11.3% for various models by regularizing only 3%–15% of their attention heads.
By Jiaqi Deng, Zonghan Wu, Zhan Heng, Xiaoshui Huang, Huan Huo, Guandong Xu
Albireo is an adaptive, energy‑efficient inference framework for video object detection on edge devices that wraps existing detectors without modification. It uses a 10‑dimensional Kalman filter per active object to decide when to skip detector calls, predicting bounding boxes on skipped frames at near‑zero GPU cost. Evaluated on BDD100K with YOLO and RF‑DETR detectors on NVIDIA Jetson AGX Thor and Orin, Albireo maintains AP@50 within ±1.2 pp of full‑frame inference while reducing energy consumption by 12.1–17.6 % and improving accuracy for some models.
By Amir Taherin, Jos\'e Cano, Bin Ren, Yanzhi Wang, David Kaeli
The paper presents an inference-only pipeline that extends the frozen NER model MahaNER‑BERT to document‑level prediction using overlapping sliding windows, eliminating the need for retraining or architectural changes. The approach is evaluated on six document‑level corpora derived from the MahaNER test set, employing two repetition strategies (Normal Repeat and Random Repeat) at three length levels and various window configurations. Results show the model maintains a macro F1‑score of up to 0.8902 with minimal variation, outperforming non‑windowed methods by avoiding boundary‑fragmentation errors and achieving more stable document‑level performance.
By Hariom Ingle, Ronit Ghode, Ishwari Gondkar, Jidnyasa Harad, Ravindra Murumkar, Raviraj Joshi
Calpric is a system that combines automatic text selection, segmentation, active learning, and crowdsourced annotation to create a large, balanced training set for privacy policy classification. By simplifying the labeling task, it enables untrained crowd workers to match the performance of trained annotators and reduces inter‑annotator disagreement, cutting labeling costs. The approach yields a dataset of 16,000 policy text segments across nine data categories and produces models that deliver accurate, fine‑grained labels at a cost of roughly $0.92–$1.71 per segment.
By Wenjun Qiu, David Lie, Lisa Austin
PTC-Bias is a two-stage framework that improves contextual biasing in speech large language models by using phoneme-level temporal competition. In the first stage, PTC Retrieval performs frame-synchronous phoneme decoding to generate a compact shortlist of bias words and their speech intervals. The second stage, PTC Correction, applies a local competition between retrieved candidates and mismatched transcript spans within those intervals, reducing near-homophone and word-segmentation errors without extra SpeechLLM passes. Experiments on LibriSpeech demonstrate consistent gains across two SpeechLLMs, with PTC-Bias reducing B-WER by up to 23.9% relative to CTC-Filter while keeping U-WER nearly unchanged.
By Zhiqi Ai, Han Cheng, Shiyi Mu, Yongjin Zhou, Shugong Xu
The paper introduces LowBridge, a method for cross‑modal medical image segmentation that leverages shared low‑level features such as edges between MRI and CT scans. It trains a generative model to reconstruct source‑modality images from edge maps and then trains a segmentation network on these generated images. At test time, edge features from target‑modality images are fed into the generative model to produce source‑style images, which are segmented by the pretrained network, achieving state‑of‑the‑art results on multiple public datasets.
By Pengfei Lyu, Pak-Hei Yeung, Jing Xia, De Hu, Xiaosheng Yu, Jianning Chi, Chengdong Wu, Jagath C. Rajapakse
The paper presents a 3D foundation model for light sheet fluorescence microscopy (LSM) that is pretrained on a large curated set of 3D images from various organisms, stains, and imaging protocols. By jointly optimizing for masked reconstruction and image‑text alignment, the model learns transferable volumetric representations that dramatically reduce the need for annotated data. The pretrained backbone enables efficient few‑shot adaptation to downstream tasks such as segmentation, classification, and deblurring, consistently outperforming baselines according to standard metrics and expert evaluation.
By Adina Scheinfeld, Haotan Zhang, Shang Mu, Rudolf L. M. van Herten, Lucas Stoffl, Ali Erturk, Zhuhao Wu, Johannes C. Paetzold
The paper introduces Smol‑VL‑BLV, a compact vision‑language model designed for blind and low‑vision users. It employs a 500M decoder transformer with teacher‑student distillation and Group Relative Policy Optimization to add spatial detail, directional cues, and hazard detection to post‑training. After a lightweight finetuning step, the model achieves significant gains on spatial, social, OCR, and VQA benchmarks while remaining under 450 MB and running entirely offline on a mid‑range Android phone.
By Rishabh Choudhary, Shreyansh Raj, Umesh Goyal, Shubh Kashyap, Shrestha Kumar, Sushovan Jena, Komal Kumar, Hisham Cholakkal, Aditya Nigam
This study evaluates the instance‑segmentation performance of YOLOv11 and YOLOv8 on immature green apples in orchard settings. YOLO11n‑seg achieved the highest mask precision (0.831), while YOLO11m‑seg and YOLO11l‑seg excelled in non‑occluded and occluded fruitlet segmentation. YOLOv8n, however, outperformed the YOLO11 series in inference speed, reaching 3.3 ms compared to 4.8 ms for the fastest YOLO11 model.
By Ranjan Sapkota, Manoj Karkee
SplatLabel is an automated pipeline that uses a 4D Gaussian representation to generate LiDAR segmentation and semantic occupancy grids with predictive confidence. It models dynamic scenes through an explicit temporal manifold, tracking moving actors without requiring pre‑annotated 3D bounding boxes. By integrating 360‑degree LiDAR depth maps and distilling soft probabilities from 2D models, it resolves semantic ambiguities over time and space, and evaluates pseudo‑labels via a selective classification framework that balances precision and recall.
By Nitya Nanvani, Andras Palffy, Holger Caesar
The paper presents a training‑free method for detecting which holds a climber uses in sport climbing videos by leveraging a frozen foundation pose model (Sapiens) that provides fingertip and toe keypoints. Using a simple proximity test, mutual exclusion, and a temporal‑persistence rule, the approach achieves high F_1 scores (up to 90.2%) on the Way Up dataset without any climbing‑specific training, outperforming repurposed pose pipelines. The resulting automatic predictions enable accurate coaching statistics, such as climb time and pace, with Pearson correlations of 1.00 and 0.94 respectively.
By Abu Bakar, Abdullah Aftab, Amir Hamza
CoSWA-YOLOv12 is a compact YOLOv12 instance‑segmentation detector designed to improve detection of tiny malaria parasites in microscopy images. It introduces a Cooperative Scale‑adaptive Wasserstein Assignment that applies a Wasserstein distance‑based label assignment inversely proportional to object size, a wavelet detail residual, and a min‑max Gaussian regression loss, all of which preserve pretrained weights. On a five‑class Rwandan thick‑smear dataset, the method raises P. falciparum recall from 0.63 to 0.74, increases mAP@50 from 0.73 to 0.81, and reduces missed detections from 38% to 15%.
By Ahmed Tahiru Issah, Carine Mukamakuza
UltraBench 2 is a new benchmark designed to evaluate vision foundation models on ultrasound images, addressing the lack of standardized tests in this area. It covers a wide range of anatomical structures and tasks, emphasizing reproducibility and ease of use. The authors compare existing models, finding that ultrasound-specific pretraining still outperforms on classification, while general-purpose models have matched performance on segmentation.
By Ashwath Radhachandran, Adam Tupper, Christian Gagn\'e, William Speier
OncoVision is a privileged‑information training framework that learns from mammography images and clinical data during training but performs inference using only mammographic images. It employs an attention‑based encoder‑decoder to jointly segment masses, calcifications, axillary findings, and breast tissue, and predicts ten structured clinical features such as BI‑RADS. Two late‑fusion strategies (Independent and Dependent) integrate imaging, radiomic, and clinical information to improve diagnostic precision, and a retrospective multi‑reader study showed higher diagnostic confidence, reduced reading time, and segmentation accuracy comparable to or better than radiologists.
By Istiak Ahmed, Galib Ahmed, K. Shahriar Sanjid, Md. Tanzim Hossain, Md. Nishan Khan, Md. Misbah Khan, Md. Arifur Rahman, Sheikh Anisul Haque, Sharmin Akhtar Rupa, Mohammed Mejbahuddin Mia, Mahmud Hasan Mostofa Kamal, Md. Mostafa Kamal Sarker, M. Monir Uddin
Background: Artificial intelligence (AI)-based glaucoma detection from colour fundus photographs (CFP) offers scalable screening, but performance may decline on external datasets because of difference...
The paper introduces EASE, an encoder‑only model for Audio‑Visual Semantic Segmentation that eliminates redundant components found in prior Transformer‑based approaches. EASE achieves state‑of‑the‑art accuracy while running at up to 365 FPS—three times faster than previous models—and can be trained in under 11 GPU‑hours. The authors provide code, weights, and samples, positioning EASE as a scalable foundation for future research and real‑time applications.
RGBD20K is a new large-scale RGB‑D dataset designed to advance semantic segmentation research. It contains 20,000 image pairs annotated with 160 fine‑grained categories, far exceeding the diversity of existing benchmarks such as NYUv2 and SUN RGB‑D. The authors also provide high‑fidelity annotations and introduce a score‑purified fusion (SPF) method that achieves state‑of‑the‑art results on multiple benchmarks.
M3D‑Net is a mammography encoder that hierarchically coordinates multi‑scale coordinate attention, bounded dynamic feature reuse, and differential attention through resolution‑aware operator placement. It preserves earlier features within stages, integrates local and global context via coordinate‑aware aggregation, and applies differential attention at coarse resolutions. In image‑only classification on AISSLab mammography and an adapted image‑clinical model on BrEaST ultrasound, M3D‑Net achieves the highest validation accuracy and lowest endpoint cross‑entropy loss compared to EdgeNeXt, RepViT, and TransXNet, with accuracies of 97.78% and 80.39% respectively.
By Zheng Yu, Xinhang Li, Jiabao Gao, Boyang Wang, Xiang Li
nnFoundation introduces complementary convolutional and transformer-based 3D foundation models for radiology, trained on 2.1 million CT, MRI, and PET volumes from 125 datasets. The models are evaluated on 108 tasks—including segmentation, detection, classification, report generation, and image retrieval—under domain shift, low-data, and low-compute scenarios, consistently outperforming prior 3D foundation models and training from scratch. Performance varies by task type, with convolutional models excelling at spatially localized tasks and transformer models at global semantic reasoning, and dynamic alignment with dataset characteristics further enhances transferability.
By Constantin Ulrich Harsy, Tassilo Wald, Karol Gotkowski, Yannick Kirchhoff, Marcel Knopp, Maximilian Rokuss, Elisa Stegmeier, Philipp Schader, Dasha Trofimova, Raphael Stock, Kim-Celine Kahl, Stephen Schaumann, Selen Erkan, David Zimmerer, Stefan Denner, Moritz Langenberg, Sebastian Ziegler, Katharina Eckstein, Maximilian Fischer, Jonathan Suprijadi, B\'alint Kov\'acs, Benjamin Hamm, Anand Deshpande, Dimitrios Bounias, Nico Disch, Shuhan Xiao, Jessica K\"achele, Jan Sellner, Rajesh Baidya, Jeremias Traub, Lars Kr\"amer, Maximilian Zenk, Tim R\"adsch, Stefan Dvoretskii, Robin Peretzke, Jonathan Deissler, Alexandra Ertl, Partha Ghosh, Kris Dreher, Stefan Dinkelacker, Annika Reinke, Evangelia Christodoulou, Numan Saeed, Yoland Savriama, Santiago Estrada, David K\"ugler, Laura Alexandra Daza Barragan, Cristina Isabel Gonzalez Osorio, Jan Peeken, Michael Baumgartner, Marvin Teichmann, Guillaume Chabin, Matthias Kirchler, Valentin Koch, for the ALFA study, Markus Hohenhaus, Dimitri Koslov, Nina Decker, Mohammad Yaqub, Arnd Heuser, Martin Reuter, Julia A. Schnabel, Tobias Heimann, Florin Ghesu, Paul Brachmann, Claus P. Heu{\ss}el, Alexander Radbruch, Gianluca Brugnara, Aditya Rastogi, Martha Foltyn-Dumitru, Heinz-Peter Schlemmer, Ignaz Reicht, Julius C. Holzschuh, Michael Bach, Bram Stieltjes, Kai Schlamp, Lena Maier-Hein, Marco Nolden, Ralf Floca, Paul F. J\"ager, Philipp Vollmuth, Fabian Isensee, Klaus H. Maier-Hein