BMASH: Ball-Motion-Aware Soccer Header Spotting
arXiv:2609.39300v1 Announce Type: new Abstract: Recent advances in computer vision have made broadcast sports videos increasingly useful for event analysis, performance assessment, and player-safety...
Detection, segmentation, depth and recognition research, plus the vision backbones that keep displacing the last generation.
arXiv:2609.39300v1 Announce Type: new Abstract: Recent advances in computer vision have made broadcast sports videos increasingly useful for event analysis, performance assessment, and player-safety...
arXiv:2609.39429v1 Announce Type: cross Abstract: Uncertainty Quantification (UQ) is a key requirement for trustworthy AI in high-stakes medical image analysis. In this work, we evaluate UQ in a mult...
arXiv:2609.40356v1 Announce Type: cross Abstract: Recent video generation is increasingly realistic and controllable, yet video editing remains less developed, particularly for precise local edits th...
arXiv:2609.39116v1 Announce Type: new Abstract: Prior-free 6D object pose tracking seeks to recover the trajectory of an unseen object from a single RGB video without object-specific CAD models, pose...
arXiv:2609.35490v2 Announce Type: replace Abstract: Generalist multitasking vision models aim to unify multiple vision tasks within a single framework, enabling more efficient and versatile learning....
arXiv:2609.38811v1 Announce Type: new Abstract: Metal additive manufacturing parts are inspected by X-ray computed tomography, where labelled data is scarce, the pores and inclusions that matter span...
arXiv:2609.38930v1 Announce Type: cross Abstract: In semantic segmentation, a recent line of RankSEG methods directly optimizes Dice/IoU scores at inference time, improving alignment with evaluation...
arXiv:2606.28444v2 Announce Type: replace-cross Abstract: Classical universal approximation theorems (UAT) establish the expressive power of sigmoidal multilayer perceptrons, but they do not specify...
arXiv:2605.27696v3 Announce Type: replace Abstract: Discrete visual tokenizers map images to ordered sequences of tokens, providing a natural representation for structural scene descriptions. Most us...
arXiv:2609.39001v1 Announce Type: cross Abstract: LLM services are billed per token and context windows are measured in tokens, yet the number of tokens needed for the same content varies across lang...
arXiv:2609.39013v1 Announce Type: new Abstract: EVICALC, our system for the DocSem shared task, achieved 8.61% joint accuracy on 1,730 tasks in the official final test evaluation. It reads a PDF, sel...
The paper investigates how raw or minimally processed satellite imagery affects onboard AI object detection for space missions. By systematically degrading Very High Resolution Maxar images in terms of Signal‑to‑Noise Ratio, Modulation Transfer Function, and Ground Sampling Distance, the authors evaluate three lightweight detectors—YOLOv5s, YOLOX‑S, and NanoDet—on the resulting data. Results show that image quality impacts detection performance in a degradation‑specific way, with GSD consistently shifting performance, while MTF and SNR effects vary by model and resolution; severe blur‑plus‑noise combinations cause the greatest losses.
The paper introduces a consensus‑aware multi‑source fusion framework for reference‑guided camouflaged object detection. It couples trainable PVTv2 query features with frozen DINOv3 representations, using reference‑conditioned correlation to select foundation‑model evidence before multi‑scale fusion. The method also aggregates multiple references via cross‑reference consensus aggregation and injects reference information at semantic depths matched to the query features, achieving complementary improvements in experiments.
The paper introduces Event-Driven Refresh + Recurrence Memory (EDRRM) to improve Referring Video Object Segmentation (RVOS). EDRRM selectively re-invokes the Sa2VA model at stable change points, using an EMA‑smoothed event score from tracking cues and a recurrence memory that retrieves anchor frames via CLIP similarity. Experiments on Ref‑DAVIS17, MeViS, and ReVOS show that EDRRM maintains or surpasses J&F scores while reducing refresh calls and false‑positive failures, with modest overhead compared to Sa2VA inference.
PCB-MC is a new dataset for detecting missing components on printed circuit boards, featuring 197 distinct board types with footprint-level annotations derived from the RF100 dataset. The dataset includes multiple augmented samples per board type and introduces board type-aware cross‑validation splits to prevent layout leakage between training and test sets. Benchmark results show that supervised models suffer high false‑negative rates on unseen board designs, while unsupervised anomaly detection methods fail due to misalignment with board-specific references, highlighting the remaining challenges in missing component detection across diverse PCB layouts.
The paper demonstrates that a stochastic hybrid system (SHS), which combines continuous dynamics governed by a stochastic differential equation (SDE) with discrete resets triggered by a Markov kernel, can be approximated by a single SDE in a higher‑dimensional latent space. By encoding the reset branches with auxiliary variables, the resets become deterministic, allowing the system’s manifold to be glued and embedded into Euclidean space. This embedding eliminates the need for explicit reset terms in the hybrid Fokker‑Planck equation, and the authors propose a loss function that matches evolving state distributions, enabling the latent SDE to recover the SHS’s probability evolution without mode labeling, trajectory segmentation, or event‑based simulations.
The paper presents an automated segmentation pipeline for whole‑slide histopathology images of colorectal cancer, labeling tumor grades 1‑3 and normal mucosa. It employs dense prediction transformers with multiple encoder backbones, overlapping patches, test‑time augmentation, and an adaptive augmentation policy guided by large language models. The approach, combined with soft‑voting ensembles and post‑processing refinements, raises the F1 score from 62.92 to 69.84 on a colorectal cancer grade dataset.
HiWE is a hierarchical world knowledge model that enables zero‑shot 3D path planning by linking visual grounding with language‑based planning through a point‑based interface. It uses PointVLM to map task‑relevant objects to image coordinates, lifts these predictions into a semantic 3D representation with depth data, and then a language planner (3DLLM) generates end‑effector waypoints and gripper commands. The system is evaluated on 14 simulated manipulation tasks and four physical‑robot tasks, with ablations on visual training data, spatial inputs, and grasp selection.
WinoTS introduces a wavelet‑based self‑distillation framework for time‑series models that uses time‑frequency augmentations to create multi‑scale structural views, avoiding distortion of signal dynamics. The method outperforms state‑of‑the‑art baselines in long‑term forecasting, cross‑domain zero‑shot transfer, and unsupervised anomaly detection, and linear probing on frozen representations often beats fully supervised training from scratch. Ablation studies show WinoTS is architecture‑agnostic and demonstrates that time‑frequency transformations offer a principled alternative to vision‑style spatial augmentations.
GroundingPI is a 4‑billion‑parameter grounding foundation model that generates points and boxes as quantized coordinates in a shared vocabulary. It is trained with multimodal and spatial pretraining, supervised fine‑tuning, and reinforcement learning, achieving a new state‑of‑the‑art average of 73.68% across 34 grounding benchmarks. As a visual backbone, GroundingPI improves performance in robotic manipulation and autonomous driving, outperforming larger models and mainstream backbones in several out‑of‑distribution settings.