Computer vision

Detection, segmentation, depth and recognition research, plus the vision backbones that keep displacing the last generation.

3,375 stories · RSS feed

arXiv AI
Sep 30

Calibrating One-Round Membership Inference with Neighbors

The paper addresses the challenge of calibrating membership inference attacks in a one‑round setting where only a single trained model is available. It proposes using neighboring data points of the target to approximate the calibration that reference models normally provide, and demonstrates that querying these neighbors—especially against early training checkpoints—enhances the membership signal. Experiments on three image classification datasets and training setups show that this neighbor‑based approach yields strong attack performance without extra training cost.

By Francesco Rita, Jie Zhang, Florian Tram\`er
arXiv Computer Vision
Sep 30

GA-EIRFS: A Geometry-Augmented Repeat-Factor Sampling Method for Long-Tailed LiDAR 3D Object Detection

GA-EIRFS is a detector‑agnostic sampling strategy that augments frequency‑based repeat‑factor sampling with a fixed geometry score derived from point count, surface‑normal entropy, and surface coverage. It only alters frame‑sampling probabilities, leaving the underlying detector and inference pipeline unchanged. Experiments on nuScenes show consistent improvements in mean average precision and the nuScenes detection score across multiple seeds and backbones, with notable gains for rare classes such as bicycles.

By Taufiq Ahmed, Constantino \'Alvarez Casado, Daniel Herrera Castro, Sasan Sharifipour, Abhishek Kumar, Miguel Bordallo L\'opez
arXiv Computer Vision
Sep 30

EviViT: Evidence-Adaptive Vision Transformers for Fine-Grained Perception

EviViT is a lightweight attachment for pretrained vision transformers that learns where to focus detail in high‑resolution images. It uses human visual‑search traces to supervise a question‑conditioned evidence density, guiding regional re‑reading and efficient visual token allocation. The method connects regional features to the global scene via a sparse, coordinate‑aware bridge, improving fine‑grained accuracy across nine host models while using fewer tokens than global‑only processing.

By Yaoxin Niu, Zhangquan Chen, Yang Zhang, Xiang An, Zhumei Wang, Chih-Ting Liao, Hongkun Cao, Ruqi Huang
arXiv AI
Sep 30

Beyond Sub-Gaussian Detector Scores: Robust Weighted Profile-Loss Change Point Detection for Human-LLM Text Segmentation

The paper introduces Robust Weighted Profile-Loss Change Point Detection (RWCP), a method for locating authorship transitions in mixed human‑LLM documents using detector scores with varying reliability. RWCP combines capped reliability weights, Huber profile gains, and a narrowest‑over‑threshold search to express the population gap as a merge cost, enabling recovery of change points without a closed‑form nonlinear center. Experiments on five benchmark families show that core RWCP reduces WindowDiff by 17.6% compared to weighted change‑point detection, and an extended variant RWCP‑R further improves performance, especially for isolated changes.

By Wan Tian, Zhongyi Li, Yawen Li, Rui Zhang, Yijie Peng, Fuzhen Zhuang
arXiv AI
Sep 30

Co-PiLOT: Constrained Physics-Informed Latent Optimization for Target-Driven Inverse Design

Co-PiLOT is a latent optimization framework that maps candidate physical structures through a generative encoder-decoder, using the decoder as a learned validity prior and performing physics-informed black-box optimization in latent space. It is applied to the inverse design of magnesium alloy microstructure/texture, employing a vision transformer encoder paired with latent diffusion, diffusion transformer, and rectified-flow transformer decoders trained on an 80,000-sample EBSD dataset to produce a minimal bottleneck representation. The MERIDIAN optimizer, driven by a deep-kernel Gaussian process and failure-aware feasibility prediction, achieves the best target-driven objective score within 160 simulations, reducing relative target error by 3–22% compared to seven baseline methods.

By Mahish K. Guru, Mayank Nagar, Ayush vyas, Jan Bohlen, Roland Aydin, Noomane Ben Khalifa
arXiv Computer Vision
Sep 30

PCaPaint: Prostate Cancer Inpainting by Mitigating Shortcut Learning

PCaPaint is a prostate cancer inpainting method that uses latent diffusion models (LDMs) and introduces a conditioning strategy where the condition image is filled with Gaussian noise to mitigate shortcut learning. It also proposes a new training objective that focuses on errors within the lesion region and a multi‑sequence latent design that separately compresses T2w and DWI&ADC scans to preserve their distinct frequency characteristics. Experiments show that the synthetic data generated by PCaPaint improves downstream tasks such as prostate lesion segmentation, patient‑level classification, and lesion‑level detection, outperforming recent state‑of‑the‑art LDM‑based tumor inpainting methods in both performance and image quality.

By Levente Lippenszky, Hongxu Yang, Marcell D\"om\"ot\"or, Krisztian Koos, L\'aszl\'o Rusk\'o
arXiv Machine Learning
Sep 30

Multi-task learning for the automatic grading of enlarged perivascular space burden using MRI

arXiv:2609.37387v1 Announce Type: cross Abstract: Enlarged perivascular spaces (PVS) visible in brain magnetic resonance imaging (MRI) are increasingly thought to be linked to poor brain health. PVS...

By Jesse Phitidis, William N. Whiteley, Joanna M. Wardlaw, Miguel O. Bernabeu, Yajun Cheng, Xiaodi Liu, Junfang Zhang, Una Clancy, Stephen Makin, Roberto Duarte Coello, Susana Mu\~noz Maniega, Mark E. Bastin, Simon R. Cox, Maria del C. Vald\'es Hern\'andez
arXiv Computer Vision
Sep 30

Sparse-View Interpretable 3D Animal Behavior Representations for Neural Encoding and Decoding

arXiv:2609.36217v1 Announce Type: new Abstract: A deeper understanding of brain function requires a precise, structured characterization of behavior.Yet, extracting behavioral representations from vi...

By Xinming Dai, Qihang Jin, Tianshu Tan, Baiyuan Chen, Hanrui Lyu, Lenny Aharon, Kyle Daruwalla, Xun Helen Hou, Matthew R. Whiteway, Liam Paninski, Yizi Zhang