Computer vision

Detection, segmentation, depth and recognition research, plus the vision backbones that keep displacing the last generation.

3,374 stories · RSS feed

arXiv Machine Learning
Oct 3

Neural scaling laws and evolution of learnable activation functions of Kolmogorov-Arnold networks

The paper studies Kolmogorov‑Arnold Networks (KANs), a neural architecture that treats activation functions as learnable components, offering improved interpretability for scientific applications. It investigates how KANs scale with dataset size on image classification tasks (MNIST, Fashion‑MNIST) and a magnetic‑parameter regression task, revealing a broken neural scaling law that transitions from a faster to a slower decay of test loss as data grows. The authors also analyze how the learned activation functions evolve from simple linear approximations to more complex, interpretable symbolic forms as more data is provided.

By Tilen Cadez, Sanghoon Lee, Kyoung-Min Kim
arXiv Machine Learning
Oct 3

Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models

The paper extends mechanistic interpretability of large language models by modeling concepts as low‑dimensional non‑linear manifolds rather than linear subspaces. It introduces a concept‑based alignment (CBA) score to compare these manifolds across layers and models, revealing block structures in intermediate layers, a shift from syntax‑dominated to mixed syntactic‑semantic concepts, and training‑dependent multilingual sharing. The study also shows that alignment patterns differ across model families and training stages, with adjacent stages aligning more closely than distant ones.

By Tido Specht, Elias Benedict Krey, Nils Neukirch, Nils Strodthoff
arXiv AI
Oct 2

A Framework for Egocentric and Exocentric Procedural Understanding via Temporal Segmentation and Semantic Abstraction

The paper introduces a compact framework that transforms continuous multimodal workplace video into a structured Procedural State Memory called a Work Environment Model (WEM). Using event segmentation theory, it detects segment boundaries based on changes in visual context, location, motion, narration, gaze/object interaction, and optional exocentric workspace evidence, then abstracts each segment into an evidence‑linked event card. These event cards incrementally update the WEM, enabling efficient, auditable documentation and retrieval while respecting on‑premise privacy constraints, and the authors evaluate the system on segmentation quality, memory compression, retrieval fidelity, and long‑horizon QA.

By Vivek Chavan, J\"org Kr\"uger
arXiv AI
Oct 2

Modeling The Object Representations Underlying Human Physical Reasoning

The study investigates how vision models represent objects compared to humans during physical reasoning tasks. By comparing segmentation outputs from several architectures (DINOv2, SegFormer, DeepLabV3+, UPerNet) to human behavior in time‑to‑collision and change‑detection tasks, the authors find that briefly trained models are too coarse, fully trained models are too fine, and an intermediate training stage best matches the coarse, volumetric bodies humans use. Larger models reach this intermediate alignment earlier, suggesting that resource constraints drive the emergence of human‑like representations in general‑purpose vision systems.

By Andrey Gizdov, Andrea Procopio, Lorenzo Caputi, Georgi I. Ivanov, Yichen Li, Daniel Harari, Tomer Ullman
arXiv Computer Vision
Oct 2

Domain generalization and synthetic data in object detection: the enabler, the probe, and the gap

The paper reviews domain generalization (DG) for object detection, highlighting its scarcity and unique challenges such as localization and multi-scale representation. It examines synthetic data from three angles: as an enabler that diversifies and aligns training data, as a probe that allows controlled experiments to uncover failure modes, and as a source of a synthetic‑to‑real gap that hampers deployment on real imagery. The authors conclude that future DG research must develop representation‑aware methods that explicitly address both localization and classification under domain shift.

By Elfi I. S. Hofmeijer, Ella P. Fokkinga, Friso G. Heslinga, Klamer Schutte, J\"orgen M. Karlholm
arXiv Computer Vision
Oct 2

Learned Suppression for 3D Keypoint Detection with a Graph-Transformer Backbone

The paper introduces a 3D keypoint detector that replaces traditional heuristic post‑processing with a learned suppression module. This module, implemented as a graph neural network over candidate keypoints, learns to keep, suppress, or relocate points, and is built on a Point Transformer backbone extended with a directional graph neural network. Experiments show the approach outperforms DBSCAN and greedy non‑maximum suppression, surpasses KeypointDETR on most KeypointNet categories, and achieves top performance on the Building3D benchmark while remaining competitive on the Tallinn split.

By Batuhan Arda Bekar, Can Sar{\i}, H\"useyin Can G\"ulkan, Bar{\i}\c{s} \"Ozcan
arXiv Computer Vision
Oct 2

Impact of Patient Orientation in Single- and Multi-View Camera Environments for AI-based Rehabilitation Monitoring

The paper introduces REHAB26-ViewAngles, a dataset of correct and incorrect rehabilitation exercise videos captured from many camera angles, and a new separability metric to evaluate pose‑estimation algorithms. It evaluates single‑camera 2D and 3D pose estimation and four multi‑camera approaches, showing that an optimally placed 2D camera can boost separability by 16.9% over a frontal view and that combining two views can improve accuracy by up to 13.1%. These findings provide concrete guidance for setting up camera systems in home and clinical rehabilitation monitoring.

By Miriama J\'ano\v{s}ov\'a, Andreas Lang, Petra Budikova, Jan Sedmidubsky
arXiv AI
Oct 2

GPart: End-to-End Isometric Fine-Tuning via Global Parameter Partitioning

GPart introduces a new parameter‑efficient fine‑tuning technique that directly maps a low‑dimensional trainable vector into the full weight space using a sparse, isometric partition matrix. Unlike LoRA, GPart eliminates the bilinear reconstruction step, preserving exact end‑to‑end isometry and reducing the checkpoint to just the vector and a random seed. Experiments across NLP, vision, and reasoning tasks show that GPart matches or surpasses existing PEFT methods while using far fewer parameters and offering a simpler, more tractable parameterization.

By Paolo Mandica, Micha{\l} Brzozowski, Zuzanna Dubanowska, Neo Christopher Chung
arXiv Machine Learning
Oct 2

Function-Space Transformer with Adaptive Anchors

The paper introduces the Function‑Space Transformer (FST), a neural framework that learns from functions by using a spatially adaptive continuous latent representation. FST places features at anchors whose positions are predicted from input samples and refines these anchors through recursive function‑space interactions, allowing the representation to adapt to each input rather than relying on a fixed grid. Experiments on PDEBench (Burgers and Darcy flow) and ImageNet‑1K show that FST outperforms or matches state‑of‑the‑art baselines such as Perceiver IO and the Vision Transformer while using fewer parameters.

By Guorui Sang, Pedram Rooshenas
arXiv Computer Vision
Oct 2

From Image Latent Space to Fuzzy Rules: Interpretable Analysis of Gastrointestinal Foundation Model

The paper introduces a prototype‑based fuzzy‑rule framework that interprets patch‑level features from pretrained foundation models without fine‑tuning. By clustering class‑specific prototypes in the feature space, the method translates latent representations into human‑readable IF‑THEN rules, achieving accuracy comparable to black‑box classifiers on gastrointestinal imaging tasks. The extracted rules are also used to analyze synthetic medical images, revealing where generators diverge from real tissue.

By Michael D. Vasilakakis (Department of Computer Science and Biomedical Informatics, University of Thessaly, Lamia, Greece), Dimitris K. Iakovidis (Department of Computer Science and Biomedical Informatics, University of Thessaly, Lamia, Greece)
arXiv Computer Vision
Oct 2

Color Independent Word Segmentation From Transcribed Bangla Passages

The paper presents a color‑independent word segmentation method for handwritten Bangla text images. It works on smartphone‑captured images regardless of paper color or ink type, and the custom dataset includes various real‑world challenges such as shadows. The system achieves 90.60 % recall, 91.80 % precision, and 91.20 % F1‑score on 7,374 words.

By Faias Satter, Noor Masrur, Sk. Md. Masudul Ahsan
arXiv Computer Vision
Oct 2

Dyna3: VLM-Guided Training-Free 4D Reconstruction via Depth Foundation Models

Dyna3 is a training‑free framework that extends the depth foundation model DA3 to perform 4D dynamic scene reconstruction without fine‑tuning. By leveraging DA3’s cross‑view features and a best‑match search, it distinguishes static surfaces from moving objects, and uses vision‑language models to generate semantic prompts for SAM 3 to achieve precise instance‑level segmentation. Experiments on four datasets show Dyna3 outperforms correspondence‑trained methods, improving dynamic object segmentation by +5.5 pp, speeding pose estimation 13×, and reducing memory usage 4–8×.

By Xinhao Xiang, Weiyang Li, Zhijie Zheng, Abhijeet Rastogi, Jiawei Zhang
arXiv Computer Vision
Oct 2

FiVOS: A Fish Segmentation Algorithm Based on Interactive Video Object Segmentation and Filter Enhancement

FiVOS is an interactive video object segmentation algorithm tailored for fish segmentation in aquaculture. It introduces a mask block filter and serial noise filters to detect and correct erroneous propagated masks early, improving accuracy and robustness. Experiments show FiVOS achieves state‑of‑the‑art performance on two newly constructed fish‑specific datasets.

By Yuqing Duan, Song Zhang, Shili Zhao, Daoliang Li, Ran Zhao
arXiv Computer Vision
Oct 2

Synthetic training for long-tail haemorrhagic lesion segmentation in data-scarce settings

The paper introduces a synthetic training framework for segmenting long‑tail haemorrhagic lesions such as cerebral microbleeds (CMBs) and cortical superficial siderosis (cSS) without requiring real lesion annotations. By starting from anatomical brain parcellations, the method applies spatial augmentation, voxel resampling, and procedural insertion of lesion labels guided by clinical priors, then synthesises images with randomized intensity, blurring, and Rician noise. Models trained on these synthetic image‑label pairs outperformed classical filter baselines, achieving higher AUPRC and AUROC for both cSS and CMB segmentation.

By Yuan Cao, Sumeet Dash, Antonia Zachariadis, Stefanie Schreiber, Katja Neumann, Jose Bernal
arXiv Computer Vision
Oct 2

PAGER: Partial-to-global Alignment via Geometric and Relational Distillation

PAGER is a label‑free adaptation method that aligns partial, viewpoint‑dependent 3D observations with a frozen global semantic space. It uses matched‑point feature alignment and relational supervision to anchor partial features to their global counterparts while preserving similarity structure, all without altering the pretrained encoder or global probe. Experiments show PAGER outperforms label‑supervised PEFT on Sonata and Concerto, and achieves superior zero‑shot transfer from ScanNet to ScanNet++ compared to fully fine‑tuned Sonata.

By Akira-Miranda Adeyomi Adeniran-Lowe, Binod Singh, Lars Arnold Dethlefsen, Lazaros Nalpantidis, Theodora Kontogianni
arXiv Computer Vision
Oct 2

Fusing Visual and Textual Representations via Multi-layer Fusing Transformers for Vietnamese Visual Question Answering

The paper introduces a Multi-layer Fusing Transformer that uses cross‑attention to merge visual and textual features across multiple layers, enabling extraction of information from low to high levels. Experiments and ablation studies on the ViVQA dataset demonstrate that this architecture outperforms competitive baselines for Vietnamese Visual Question Answering. The work addresses a notable gap in VQA research for Vietnamese, a language with limited prior resources.

By Cong Phu Nguyen, Huy Tien Nguyen, Tung Le
arXiv Computer Vision
Oct 2

GenCOPE: Syn2Real Generalized Category-Level Object Pose Estimation for Robotic Picking

GenCOPE introduces a synthetic-to-real (Syn2Real) approach for category-level object pose estimation (COPE) that eliminates the need for labor-intensive real-world data collection. By learning domain-invariant representations through 2D and 3D semantic consistency constraints and employing an end-to-end pose regression framework with 2D-3D cross consistency, the model achieves robust generalization across synthetic and real domains. The architecture relies solely on global features, resulting in a lightweight and efficient design validated on REAL275, Wild6D, and real-world robotic manipulation scenes.

By Jian Liu, Wei Sun, Zhenqi Dai, Hui Yang, Jian Xiao, Nicu Sebe, Na Zhao
arXiv Computer Vision
Oct 2

Signal-Noise Factorization Isolates Nuisance Variation into Removable Subspaces

The paper introduces regularizers that enforce signal‑noise factorization (SNF) and signal‑signal factorization (SSF) during training of deep neural networks. Experiments on CIFAR‑100 show that SNF regularization improves classification accuracy, while SSF does not. On the BloodMNIST dataset with varying corruption levels, SNF yields even larger gains, and analysis reveals that SNF isolates noise into distinct subspaces, enabling projection of corruption‑induced directions and further accuracy improvements.

By Sakin Kirti, Joel Zylberberg
arXiv Computer Vision
Oct 2

VITA: A Multi-Source Vicinal Transfer Augmentation Method for Out-of-Distribution Generalization

The paper introduces VITA, a multi‑source vicinal transfer augmentation method designed to improve out‑of‑distribution generalization in computer vision. VITA combines tangent transfer, which generates initial augmented samples to enhance robustness against various image corruptions, with an integration step that uses a generative model to produce on‑manifold samples from multiple vicinal sources. Experiments on corruption benchmarks show that VITA outperforms existing state‑of‑the‑art augmentation techniques.

By Minghui Chen, Cheng Wen, Feng Zheng, Fengxiang He, Ling Shao
arXiv Computer Vision
Oct 2

Augmented Equivariant Mesh Networks for Anatomical Segmentation

The paper introduces EAMS, an Equivariant Anatomical Mesh Segmentor that operates directly on irregular surface geometry and remains robust to coordinate pose changes across meshes of varying resolution. Built on Equivariant Mesh Neural Networks (EMNN), EAMS combines intrinsic mesh descriptors with anatomy-aware priors, such as PCA-derived frames for dental arches and liver surfaces, and augments message passing to provide lightweight global context. Evaluated across four datasets and three clinical application areas, EAMS variants perform competitively on unperturbed inputs while maintaining stability under geometric perturbations, demonstrating that a lightweight (<2 M parameters) equivariant framework can achieve robust anatomical mesh segmentation across diverse label types.

By Daniel Saragih