The paper introduces FROST, an online framework that filters synthetic training data by estimating its utility through gradient feedback anchored in real data. FROST calibrates batch utility against recent history to decide when to filter, removing 20–30% of synthetic samples while improving performance on image classification and LLM fine-tuning tasks. The method is also applied to a large‑scale industrial ads re‑ranking system, yielding significant gains over an optimized production baseline.
By Yanran Wu, Sana Lakdawala, Renzo Tassara Miller, Chongyang Bai, Sharath Ciddu, Shivendra Pratap Singh, Kungang Li, Sandeep Pandey, Chunwei Liu
Multimodal Routing and Region Refinement for Language-Guided Medical Image Segmentation (MRSeg) is a parameter‑efficient framework that uses frozen ConvNeXt‑Tiny and PubMedBERT encoders to extract multiscale visual features and clinical text tokens. A joint router predicts a sparse mixture over low‑rank adapter bases, enabling separate adaptation for two visual scales and text while keeping feature‑specific parameters distinct. Region Bridge aggregates dense visual tokens into latent regions using text‑derived queries, refines them via self‑attention and text cross‑attention, and redistributes the refined information back to the feature maps, culminating in a multiscale decoder that combines refined semantic features with shallow image evidence. MRSeg achieves state‑of‑the‑art Dice/mIoU scores on QaTa‑COV19 and MosMedData+ with only 7.11 M trainable parameters and 7.60 GFLOPs.
By Md Maklachur Rahman, Md Hasan Al Banna, Saraf Anjum, Assame Arnob, Tracy Hammond
MLPerf Automotive is the first standardized public performance benchmark for evaluating machine learning systems used in automotive AI acceleration. Developed by MLCommons, it addresses the unique constraints of automotive workloads—such as sensor suites, safety, and real‑time processing—that existing benchmarks cannot handle. The benchmark covers perception tasks (2D/3D object detection, semantic segmentation), end‑to‑end driving, and infotainment, and includes curated models, methodology, and reference implementations available on GitHub.
By Radoyeh Shojaei, Predrag Djurdjevic, Mostafa El-Khamy, James Goel, Kasper Mecklenburg, P{\i}nar Muyan-\"Oz\c{c}elik, John Owens, Tom St. John, Jinho Suh, Arjun Suresh
Token Clustering and Semantic Sequence Mamba (STMamba) is a new approach for hyperspectral image classification that organizes sparse tokens into semantically coherent sequences. It uses a hierarchical encoder-decoder with a Token Clustering Module (TCM) to select semantic tokens and a Cross-scale Neighborhood Attention (CNA) Upsampler to restore dense features. At the micro level, density-aware clustering and a quadtree-based dynamic selection keep sparse, spatially distributed tokens, while Spatial and Spectral Semantic-wise Sequencing Mamba (SWSM) modules capture long-range spatial and spectral dependencies within homogeneous semantic token sequences. Experiments on three large-scale benchmark datasets show that STMamba outperforms state‑of‑the‑art methods in both quantitative and qualitative metrics.
By Yimin Zhu, Mahmood Elahi, Lincoln Linlin Xu
PePESeg3D introduces perception priors into a multi‑scale 3D Gaussian segmentation pipeline, integrating monocular depth and mask constraints during geometry reconstruction and dense depth‑color cues with view‑consistent centroid supervision during contrastive feature learning. This dual‑stage approach aligns geometry with semantic structure and compensates for incomplete mask supervision from 2D foundation models. Experiments on SPIn‑NeRF, LERF‑Mask, and NVOS benchmarks show state‑of‑the‑art performance in both multi‑scale segmentation and scene reconstruction.
By Sungjae Choi, Seunghee Koh, Junmo Kim
MEVL-STP introduces a two‑stage pipeline for spotting arbitrarily shaped scene text. The detection stage fuses features from six frozen vision encoders via a hierarchical Feature Pyramid Network and a Progressive Scale Expansion network to produce precise polygon masks. The recognition stage then crops these masks and feeds them to a fine‑tuned Qwen3‑VL‑8B‑Instruct model, achieving state‑of‑the‑art detection and end‑to‑end performance on CTW1500, Total‑Text, and ICDAR 2015 without synthetic pretraining.
By Aman Anand, Partha Pratim Roy, Shivakumara Palaiahnakote
The paper introduces MK‑FSS, a few‑shot segmentation framework that leverages Multimodal Large Language Models (MLLMs) to extract spatial and semantic target knowledge from query images. Spatial knowledge is encoded into a memory representation and fused with support‑guided memory via a dual‑memory debate‑fusion module, while semantic knowledge is turned into a textual feature and combined with multi‑scale query features through a progressive cross‑modal prompt generator. Together, these components produce a robust target representation that improves segmentation performance over existing methods.
By Yijun Hu, Heng Fan, Libo Zhang
MoVISA introduces Multi-Token Reasoning for Video Object Segmentation, using multiple segmentation tokens (e.g., SEG0, SEG1) instead of a single token to better localize multiple objects over time. This approach enhances fine-grained alignment between language prompts and spatio-temporal mask predictions, leading to improved performance and interpretability. On benchmarks such as MeViS, DAVIS17, ReVOS, and Ref-Youtube-VOS, MoVISA achieves significant gains, notably a 13.2% J and F improvement on MeViS and an 8.4% J and F improvement on ReVOS.
By Ruining Zhao, Ho Kei Cheng, Alexander G Schwing
The paper introduces a tool‑augmented framework that enhances a small Vision‑Language Model (Qwen3.5‑4B) with geometric tools—3D object detection, metric depth estimation, and deterministic solvers for distance, size, and bearing—to improve metric spatial reasoning. By moving metric computation from the model’s weights into explicit solvers, the approach achieves significant gains on ReVSI‑Bench tasks, notably increasing absolute distance accuracy from 0.46 to 0.74 MRA and relative direction accuracy from 25.9% to 73.4%. The modular design allows swapping in different detectors, enabling a clear separation between perception and reasoning errors, and the model can autonomously sequence the tools to match a scripted pipeline on most tasks.
By Kai Glantz, Clemens Grange
RFlash is a physics‑informed post‑processing technique that removes acoustic shadows in ultrasound by decomposing beamformed images into attenuation and scatter‑intensity maps using a differentiable radiance‑field model. It re‑renders images to eliminate depth‑dependent signal loss, effectively simulating a virtual transducer advance. Across thousands of fetal brain, abdominal, and liver scans, RFlash outperforms traditional attenuation correction, reduces prediction error by 40% in fetal brain imaging, and provides shadow‑confidence maps that enhance bone‑shadow segmentation.
By Valentin Bacher (Oxford Machine Learning in NeuroImaging Lab, University of Oxford, United Kingdom), Pak Hei Yeung (Oxford Machine Learning in NeuroImaging Lab, University of Oxford, United Kingdom, Quantitative Healthcare Analysis), Bernhard Kainz (Friedrich-Alexander-Universit\"at Erlangen-N\"urnberg, Germany, Imperial College London, United Kingdom), Madeleine K. Wyburd (Oxford Machine Learning in NeuroImaging Lab, University of Oxford, United Kingdom, Department of Computer Science, University of Copenhagen, Denmark), Nicola K. Dinsdale (Oxford Machine Learning in NeuroImaging Lab, University of Oxford, United Kingdom), Michael Gray (Institute of Biomedical Engineering, University of Oxford, United Kingdom), Ana I. L. Namburete (Oxford Machine Learning in NeuroImaging Lab, University of Oxford, United Kingdom)
SpectralCTGaussians introduces a novel spectral CT reconstruction and basis material decomposition technique that employs 3D Gaussian splatting with per‑Gaussian material fractions and energy‑dependent basis functions. By jointly optimizing these parameters across all energy channels via a differentiable polychromatic forward model, the method achieves superior novel view synthesis and higher PSNR for spectral CT volume reconstruction compared to traditional and learning‑based baselines. It also provides one‑step material decomposition with direct RGB segmentation and recovers the photoelectric basis more accurately than existing pipelines.
By Reinout Vos, Saptarshi Neil Sinha, Michael Weinmann
TopoFuse introduces a topology-aware tri-planar fusion method for 3D cryo-electron tomography segmentation. It replaces traditional loss penalties with a differentiable projection operator that identifies and sparsely edits critical voxels to enforce specified topological constraints. The approach achieves a 54% reduction in Betti number error, a 4.6-point Dice improvement, and edits only 3.1% of voxels across three benchmarks.
By Rohit Kumar Salla, Neelesh Gupta, Xingjian Li, Min Xu
The paper introduces a mesh-guided post‑processing framework that repairs broken vessel segmentations produced by nnU‑Net. By fitting a deformable template mesh to each predicted binary mask, the method conservatively reconnects disconnected components using thin bridge candidates while respecting foreground‑growth constraints. Evaluation on aorta, Circle of Willis, and pulmonary artery datasets shows that connectivity metrics (ccDice and Betti‑0) improve dramatically with negligible change in overall overlap (Dice).
By Gniewosz Drwiega, Wojciech Szymanski, Marek Wodzinski
The paper introduces SphereTrust, a method that uses frozen self‑supervised hyperspherical features to evaluate and rank candidate masks produced by foundation segmenters like SAM. By measuring angular contrast, foreground coverage, and image‑frame contact, SphereTrust can select high‑quality masks in 0.55 s per image and outperforms existing baselines on multiple segmentation tasks. The selected masks are then used as priors to train student models, improving performance on several benchmark datasets.
By Xinge Guo, Fengyang Xiao, Dingming Zhang, Yuhan Chen, Rihan Zhang, Xingjian Li, Tianyang Wang, Chunming He, Sina Farsiu
MDE-VIO integrates learned depth priors into the VINS-Mono optimization backend to improve visual‑inertial odometry in low‑texture environments. The framework enforces affine‑invariant depth consistency and pairwise ordinal constraints while filtering unstable artifacts with variance‑based gating, keeping computation within edge‑device limits. Experiments on TartanGround and M3ED datasets show the method prevents divergence and reduces Absolute Trajectory Error by up to 28.3%.
By Arda Alniak, Sinan Kalkan, Mustafa Mert Ankarali, Afsar Saranli, Abdullah Aydin Alatan
COMPASS is a completion-and-fusion framework designed for multimodal human activity recognition (HAR) and human pose estimation (HPE) when some modalities are missing at inference. It assigns each modality to a fixed slot, filling it with either an observed representation or a completion inferred from available inputs, and uses fusion‑matched supervision to train completions against real targets at the readout level. Experiments on XRF55 and MM‑Fi datasets show that COMPASS outperforms strong baselines and alternative matching strategies for both HAR and HPE.
By Hao Wang, Yanyu Qian, Pengcheng Weng, Zixuan Xia, William Dan, Yangxin Xu, Fei Wang
ConPro introduces a self‑supervised pretraining method for vessel segmentation in digital subtraction angiography (DSA) by using a contrast projection target— the normalized drop of each pixel below its temporal median. On the DIAS and DSCA datasets, ConPro outperforms training from scratch across 10%, 20%, and 50% labeled data, and it is the best among compared methods on DSCA at 20% and 50% labels. When combined with semi‑supervised training, ConPro‑derived weights boost the UniMatch baseline by 0.5–2.0 Dice points and 0.9–2.3 clDice points, achieving 75.4 Dice on DIAS and 81.3 on DSCA.
By Xinge Guo, Yuanhao Wang, Liqi Shu, Yang Liu, Min Xu
The paper presents a smartphone-based system that uses computer vision to automatically estimate vehicle speed and identify vehicles by license plate, make/model, and color. Experiments on a Brazilian dataset and real-world recordings in Austin, Texas show moderate recognition rates: 46% for license plates, 60.8% for color, 48.6% for make, and 16.89% for make/model. The study also discusses legal, technological, and practical considerations for deploying such smartphone recordings in traffic enforcement.
By Keya Li, Jahnavi Malagavalli, Lamha Goel, Tong Wang, Kara M. Kockelman
The paper "Less is More: Encoder-only Audio-Visual Segmentation" introduces EASE, an encoder-only model for Audio‑Visual Semantic Segmentation (AVSS). EASE achieves state‑of‑the‑art accuracy while running at up to 365 FPS—about three times faster than previous Transformer‑based AVSS models—and trains in under 11 GPU‑hours. The authors demonstrate that simpler, faster architectures can match or exceed the performance of more complex models across various backbones and resolutions.
By Ilpo Viertola, Vladimir Iashin, Sophie T\"otterstr\"om, Esa Rahtu
GeoRefer-Bench is a new benchmark for verifiable geospatial referring segmentation that evaluates whether models correctly resolve spatial relations in overhead imagery. Each query is expressed as an executable logical form over a metric scene graph, and predictions are scored with Exact Query Success (EQS), requiring an exact match to the query’s referent set. The dataset contains 700 UAV scenes, 26,217 instances, 142,796 spatial relations, 20,916 executable queries across five reasoning levels, and additional paraphrases, unanswerable queries, counterfactual pairs, and leakage‑controlled splits.
By Shuaishuai Cao, Min Huang, Meng Tang, Xuan Liu, Youjin Wang, Hui Lin