The paper studies Kolmogorov‑Arnold Networks (KANs), a neural architecture that treats activation functions as learnable components, offering improved interpretability for scientific applications. It investigates how KANs scale with dataset size on image classification tasks (MNIST, Fashion‑MNIST) and a magnetic‑parameter regression task, revealing a broken neural scaling law that transitions from a faster to a slower decay of test loss as data grows. The authors also analyze how the learned activation functions evolve from simple linear approximations to more complex, interpretable symbolic forms as more data is provided.
By Tilen Cadez, Sanghoon Lee, Kyoung-Min Kim
The paper extends mechanistic interpretability of large language models by modeling concepts as low‑dimensional non‑linear manifolds rather than linear subspaces. It introduces a concept‑based alignment (CBA) score to compare these manifolds across layers and models, revealing block structures in intermediate layers, a shift from syntax‑dominated to mixed syntactic‑semantic concepts, and training‑dependent multilingual sharing. The study also shows that alignment patterns differ across model families and training stages, with adjacent stages aligning more closely than distant ones.
By Tido Specht, Elias Benedict Krey, Nils Neukirch, Nils Strodthoff
The paper introduces a compact framework that transforms continuous multimodal workplace video into a structured Procedural State Memory called a Work Environment Model (WEM). Using event segmentation theory, it detects segment boundaries based on changes in visual context, location, motion, narration, gaze/object interaction, and optional exocentric workspace evidence, then abstracts each segment into an evidence‑linked event card. These event cards incrementally update the WEM, enabling efficient, auditable documentation and retrieval while respecting on‑premise privacy constraints, and the authors evaluate the system on segmentation quality, memory compression, retrieval fidelity, and long‑horizon QA.
By Vivek Chavan, J\"org Kr\"uger
The study investigates how vision models represent objects compared to humans during physical reasoning tasks. By comparing segmentation outputs from several architectures (DINOv2, SegFormer, DeepLabV3+, UPerNet) to human behavior in time‑to‑collision and change‑detection tasks, the authors find that briefly trained models are too coarse, fully trained models are too fine, and an intermediate training stage best matches the coarse, volumetric bodies humans use. Larger models reach this intermediate alignment earlier, suggesting that resource constraints drive the emergence of human‑like representations in general‑purpose vision systems.
By Andrey Gizdov, Andrea Procopio, Lorenzo Caputi, Georgi I. Ivanov, Yichen Li, Daniel Harari, Tomer Ullman
The paper reviews domain generalization (DG) for object detection, highlighting its scarcity and unique challenges such as localization and multi-scale representation. It examines synthetic data from three angles: as an enabler that diversifies and aligns training data, as a probe that allows controlled experiments to uncover failure modes, and as a source of a synthetic‑to‑real gap that hampers deployment on real imagery. The authors conclude that future DG research must develop representation‑aware methods that explicitly address both localization and classification under domain shift.
By Elfi I. S. Hofmeijer, Ella P. Fokkinga, Friso G. Heslinga, Klamer Schutte, J\"orgen M. Karlholm
The paper introduces a 3D keypoint detector that replaces traditional heuristic post‑processing with a learned suppression module. This module, implemented as a graph neural network over candidate keypoints, learns to keep, suppress, or relocate points, and is built on a Point Transformer backbone extended with a directional graph neural network. Experiments show the approach outperforms DBSCAN and greedy non‑maximum suppression, surpasses KeypointDETR on most KeypointNet categories, and achieves top performance on the Building3D benchmark while remaining competitive on the Tallinn split.
By Batuhan Arda Bekar, Can Sar{\i}, H\"useyin Can G\"ulkan, Bar{\i}\c{s} \"Ozcan
The paper introduces REHAB26-ViewAngles, a dataset of correct and incorrect rehabilitation exercise videos captured from many camera angles, and a new separability metric to evaluate pose‑estimation algorithms. It evaluates single‑camera 2D and 3D pose estimation and four multi‑camera approaches, showing that an optimally placed 2D camera can boost separability by 16.9% over a frontal view and that combining two views can improve accuracy by up to 13.1%. These findings provide concrete guidance for setting up camera systems in home and clinical rehabilitation monitoring.
By Miriama J\'ano\v{s}ov\'a, Andreas Lang, Petra Budikova, Jan Sedmidubsky
GPart introduces a new parameter‑efficient fine‑tuning technique that directly maps a low‑dimensional trainable vector into the full weight space using a sparse, isometric partition matrix. Unlike LoRA, GPart eliminates the bilinear reconstruction step, preserving exact end‑to‑end isometry and reducing the checkpoint to just the vector and a random seed. Experiments across NLP, vision, and reasoning tasks show that GPart matches or surpasses existing PEFT methods while using far fewer parameters and offering a simpler, more tractable parameterization.
By Paolo Mandica, Micha{\l} Brzozowski, Zuzanna Dubanowska, Neo Christopher Chung
The paper introduces the Function‑Space Transformer (FST), a neural framework that learns from functions by using a spatially adaptive continuous latent representation. FST places features at anchors whose positions are predicted from input samples and refines these anchors through recursive function‑space interactions, allowing the representation to adapt to each input rather than relying on a fixed grid. Experiments on PDEBench (Burgers and Darcy flow) and ImageNet‑1K show that FST outperforms or matches state‑of‑the‑art baselines such as Perceiver IO and the Vision Transformer while using fewer parameters.
By Guorui Sang, Pedram Rooshenas
The paper introduces a prototype‑based fuzzy‑rule framework that interprets patch‑level features from pretrained foundation models without fine‑tuning. By clustering class‑specific prototypes in the feature space, the method translates latent representations into human‑readable IF‑THEN rules, achieving accuracy comparable to black‑box classifiers on gastrointestinal imaging tasks. The extracted rules are also used to analyze synthetic medical images, revealing where generators diverge from real tissue.
By Michael D. Vasilakakis (Department of Computer Science and Biomedical Informatics, University of Thessaly, Lamia, Greece), Dimitris K. Iakovidis (Department of Computer Science and Biomedical Informatics, University of Thessaly, Lamia, Greece)
The paper presents a color‑independent word segmentation method for handwritten Bangla text images. It works on smartphone‑captured images regardless of paper color or ink type, and the custom dataset includes various real‑world challenges such as shadows. The system achieves 90.60 % recall, 91.80 % precision, and 91.20 % F1‑score on 7,374 words.
By Faias Satter, Noor Masrur, Sk. Md. Masudul Ahsan
Dyna3 is a training‑free framework that extends the depth foundation model DA3 to perform 4D dynamic scene reconstruction without fine‑tuning. By leveraging DA3’s cross‑view features and a best‑match search, it distinguishes static surfaces from moving objects, and uses vision‑language models to generate semantic prompts for SAM 3 to achieve precise instance‑level segmentation. Experiments on four datasets show Dyna3 outperforms correspondence‑trained methods, improving dynamic object segmentation by +5.5 pp, speeding pose estimation 13×, and reducing memory usage 4–8×.
By Xinhao Xiang, Weiyang Li, Zhijie Zheng, Abhijeet Rastogi, Jiawei Zhang
FiVOS is an interactive video object segmentation algorithm tailored for fish segmentation in aquaculture. It introduces a mask block filter and serial noise filters to detect and correct erroneous propagated masks early, improving accuracy and robustness. Experiments show FiVOS achieves state‑of‑the‑art performance on two newly constructed fish‑specific datasets.
By Yuqing Duan, Song Zhang, Shili Zhao, Daoliang Li, Ran Zhao
The paper introduces a synthetic training framework for segmenting long‑tail haemorrhagic lesions such as cerebral microbleeds (CMBs) and cortical superficial siderosis (cSS) without requiring real lesion annotations. By starting from anatomical brain parcellations, the method applies spatial augmentation, voxel resampling, and procedural insertion of lesion labels guided by clinical priors, then synthesises images with randomized intensity, blurring, and Rician noise. Models trained on these synthetic image‑label pairs outperformed classical filter baselines, achieving higher AUPRC and AUROC for both cSS and CMB segmentation.
By Yuan Cao, Sumeet Dash, Antonia Zachariadis, Stefanie Schreiber, Katja Neumann, Jose Bernal
PAGER is a label‑free adaptation method that aligns partial, viewpoint‑dependent 3D observations with a frozen global semantic space. It uses matched‑point feature alignment and relational supervision to anchor partial features to their global counterparts while preserving similarity structure, all without altering the pretrained encoder or global probe. Experiments show PAGER outperforms label‑supervised PEFT on Sonata and Concerto, and achieves superior zero‑shot transfer from ScanNet to ScanNet++ compared to fully fine‑tuned Sonata.
By Akira-Miranda Adeyomi Adeniran-Lowe, Binod Singh, Lars Arnold Dethlefsen, Lazaros Nalpantidis, Theodora Kontogianni
The paper introduces a Multi-layer Fusing Transformer that uses cross‑attention to merge visual and textual features across multiple layers, enabling extraction of information from low to high levels. Experiments and ablation studies on the ViVQA dataset demonstrate that this architecture outperforms competitive baselines for Vietnamese Visual Question Answering. The work addresses a notable gap in VQA research for Vietnamese, a language with limited prior resources.
By Cong Phu Nguyen, Huy Tien Nguyen, Tung Le
GenCOPE introduces a synthetic-to-real (Syn2Real) approach for category-level object pose estimation (COPE) that eliminates the need for labor-intensive real-world data collection. By learning domain-invariant representations through 2D and 3D semantic consistency constraints and employing an end-to-end pose regression framework with 2D-3D cross consistency, the model achieves robust generalization across synthetic and real domains. The architecture relies solely on global features, resulting in a lightweight and efficient design validated on REAL275, Wild6D, and real-world robotic manipulation scenes.
By Jian Liu, Wei Sun, Zhenqi Dai, Hui Yang, Jian Xiao, Nicu Sebe, Na Zhao
The paper introduces regularizers that enforce signal‑noise factorization (SNF) and signal‑signal factorization (SSF) during training of deep neural networks. Experiments on CIFAR‑100 show that SNF regularization improves classification accuracy, while SSF does not. On the BloodMNIST dataset with varying corruption levels, SNF yields even larger gains, and analysis reveals that SNF isolates noise into distinct subspaces, enabling projection of corruption‑induced directions and further accuracy improvements.
By Sakin Kirti, Joel Zylberberg
The paper introduces VITA, a multi‑source vicinal transfer augmentation method designed to improve out‑of‑distribution generalization in computer vision. VITA combines tangent transfer, which generates initial augmented samples to enhance robustness against various image corruptions, with an integration step that uses a generative model to produce on‑manifold samples from multiple vicinal sources. Experiments on corruption benchmarks show that VITA outperforms existing state‑of‑the‑art augmentation techniques.
By Minghui Chen, Cheng Wen, Feng Zheng, Fengxiang He, Ling Shao
The paper introduces EAMS, an Equivariant Anatomical Mesh Segmentor that operates directly on irregular surface geometry and remains robust to coordinate pose changes across meshes of varying resolution. Built on Equivariant Mesh Neural Networks (EMNN), EAMS combines intrinsic mesh descriptors with anatomy-aware priors, such as PCA-derived frames for dental arches and liver surfaces, and augments message passing to provide lightweight global context. Evaluated across four datasets and three clinical application areas, EAMS variants perform competitively on unperturbed inputs while maintaining stability under geometric perturbations, demonstrating that a lightweight (<2 M parameters) equivariant framework can achieve robust anatomical mesh segmentation across diverse label types.
By Daniel Saragih