arXiv Machine Learning

The Decision Value of Perception Compute

arXiv Computer Vision
Sep 7

Cross-Domain Tracker Adaptation Without Target-Domain Labels via Vision-Language Agents

The paper introduces a Vision‑Language Model (VLM) that acts as a diagnostic agent to adapt a detect‑to‑track system to new domains without target‑domain labels. By inspecting rendered tracking outputs, the VLM identifies failure modes and iteratively recommends parameter updates, recovering a significant portion of performance lost when transferring hyperparameters from a source domain. Experiments on MOT17→MOT20 show the VLM tuner restores 67.8% of the lost headroom, while Bayesian optimization with proxy objectives performs poorly under large domain shifts.

By Daniel Davila, Ravikumar Balakrishnan, Mike Cochran
arXiv AI
Aug 21

Active Spiking Perception: The Membrane Potential as a Belief State for Anytime 3D Point Cloud Recognition

arXiv:2608. 19232v1 Announce Type: cross Abstract: Spiking point cloud networks usually scan space in a fixed, input-agnostic order, which leaves the most distinctive resource of spiking computation, the temporal evolution of the membrane potential, unused as a locus of decision-making.

By Akarsh Jain, Arya Pawa, Ayush Debnath, Smera Rawal, Sayeed Shafayet Chowdhury
arXiv AI
Sep 3

Modeling What Changes: Sparse, Residual World Models for Object-Centric Manipulation

The paper introduces a sparse, residual world model that focuses on predicting only the changes in a scene by using a per-object change gate and a residual delta head. On a MuJoCo tabletop pushing benchmark, this approach outperforms a dense multilayer perceptron, achieving 2.5 to 4.6 times better next‑state pose accuracy with 8.6 to 11.1 times fewer parameters, maintaining high change‑detection F1 scores, and showing strong transfer across object counts. In autoregressive rollout and sampling‑based planning, the sparse model accumulates less error and enables successful planning where dense models fail.

By Param Thakkar, Parsika Paresh Shah, Manisha Sushant Gote
arXiv Machine Learning
6d ago

Decoupled Early Exits for Task-Dependent Compute Allocation in Flow-Matching VLAs

The paper introduces a framework for Flow‑Matching Vision‑Language‑Action (VLA) models that allows independent adjustment of backbone depth, action expert depth, and denoising steps. Lightweight Exit Transformers are added at intermediate layers to enable early exits, and a KV Cache synthesis mechanism manages skipped layers so the action expert can exit deeper than the backbone. Experiments on SmolVLA and π0.5 across LIBERO and Meta‑World show that joint tuning of these compute axes reduces latency by 79.2 % and FLOPs by 31.8 %, while improving mean success rate by 5.6 %.

By Riccardo Andrea Izzo, Rimvydas Rubavicius, Gianluca Bardaro, Subramanian Ramamoorthy, Matteo Matteucci, Alessandro Suglia
arXiv AI
Sep 15

Sampling headroom is not selection gain: a compute-value audit of test-time scaling for video world models

The paper introduces the Compute-Value Audit (CVA), a sequential framework that evaluates whether extra sampling during test‑time scaling for video world models actually yields a net benefit after accounting for the compute cost of generation and verification. On 192 Physics‑IQ scenes, increasing the sample pool from 4 to 16 candidates improves oracle quality by +9.23 IQ, yet common metrics such as Flow, Cycle, and VideoReward fail to reliably recover this headroom, and adaptive‑depth policies recover only 42‑69% of the potential gain. Only a few specific interventions—anchor‑explorer in a sparse PRM800K setting, MMLU‑Pro exposing a predictive‑state gap, and a privileged paired‑future upper bound—successfully pass all CVA stages, indicating that sampling headroom is valuable only when it can be converted into a reliable decision that survives the full compute charge.

By Yuhua Jiang, Junjie Lu, Feifei Gao
arXiv Machine Learning
Sep 22

Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone

The paper introduces Neural Spectral Capacity (NSC), a closed‑form metric derived from the singular‑value spectrum of weight matrices that can be computed solely from a network’s architectural specification. Unlike traditional measures such as #Params and #FLOPs, NSC captures architectural structure (depth, width, head, FFN allocations) and can be evaluated without instantiating the model, data, or gradients. Using a dynamic‑programming solver (NSC‑DP), the authors demonstrate that NSC can efficiently identify architectures that outperform existing training‑free proxies across Transformer and CNN families, and achieve state‑of‑the‑art results in tasks such as WikiText‑103 and commonsense reasoning with LLaMA‑7B. whyItMatters":"NSC provides a fast, architecture‑only proxy that outperforms conventional metrics and training‑free proxies, enabling more effective design and pruning of large models without costly training or data."

By Chenyu Zhu, Ruoyu Zhao, Zhichao Lu
arXiv Computer Vision
Aug 24

When does fusing hand-crafted knowledge with learned representations pay? A cost-normalized benchmark of stacking, substitution, and interference

arXiv:2608.21098v1 Announce Type: new Abstract: Fusing prior knowledge with data-driven learning is attractive where data is scarce, yet no controlled account says when it helps, is redundant, or har...

By Ahmad AlMughrabi, Albert Clop, Benjamin Busam, Ricardo Marques, Petia Radeva