arXiv Computer Vision By Changyi Li, Yu Xiao

Reachability Is Not Generalization: Understanding Verb--Noun Decomposition in Assembly Action Recognition

Read the original on arXiv Computer Vision →

The paper investigates how verb–noun decomposition, a common strategy for recognizing assembly actions, generalizes to novel combinations of familiar components. Through a systematic study on three datasets (MECCANO, HAViD, and IMPACT), the authors find that while decomposition avoids the zero‑probability ceiling of atomic classifiers, its performance still heavily depends on the co‑occurrence patterns seen during training. The analysis reveals that errors concentrate on the larger‑vocabulary component, that shared‑encoder training can entangle components and worsen generalization, and that these issues stem from primitive support, vocabulary asymmetry, and component entanglement.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
6d ago

Diagnosing the Sources of Compositional Failure in Vision-Language Models: A Controlled Analysis

arXiv:2609.31456v1 Announce Type: new Abstract: Vision-language models (VLMs) often struggle with compositional reasoning tasks, but the reasons for this underperformance remain unclear. A common hyp...

By Mona Gandhi, Cenk Merih Olcay, Kuan-Chieh Lo, Santiago Castro, Christopher W. Myers, Srinivasan Parthasarathy
arXiv Machine Learning
Sep 22

Comparing Latent Concept Formation in State Space Models and Transformers via Sparse Autoencoders

The paper compares latent representations in Selective State Space Models (SSMs) like Mamba and Transformers such as Pythia using Sparse Autoencoders. Across a 10‑million token corpus, 99.98% of Mamba features align closely with Pythia’s, supporting the Universality Hypothesis that core semantic representations are similar across architectures. A tiny 0.02% of features diverge, with Mamba’s recurrent bottleneck causing it to compress syntactic anomalies into polysemantic neurons, whereas Pythia’s attention can isolate distinct formatting edge‑cases.

By Rithin Nagaraj, Rupa Laalasa Oruganti, Prerna Subhashchandra Kunder, Ashwini M Joshi
arXiv Machine Learning
Sep 21

Exemplar Partitioning for Mechanistic Interpretability

The paper introduces Exemplar Partitioning (EP), an unsupervised technique that constructs interpretable feature dictionaries from large language model activations by clustering streamed activations into Voronoi regions defined by exemplars and their averages. EP allows comparison of dictionaries across layers, checkpoints, and architectures, and demonstrates utility in interpreting model behavior, tracking training dynamics, detecting hidden concepts, and enabling targeted interventions. Experiments on Gemma‑2‑2B and Llama‑3.1‑8B show EP can reveal how instruction tuning reorganizes harmful prompt activations, facilitate interventions that alter model responses, and achieve high concept‑detection performance while requiring far fewer construction tokens than comparable methods.

By Jessica Rumbelow