arXiv Machine Learning

MAxBench: A Multinomial Concept Recovery Benchmark

MAxBench is a geometry‑agnostic benchmark for evaluating how well language models recover multinomial concept representations. The study compares ten localization methods across five geometry types, six concepts, and four models, finding that affine subspaces generally steer more reliably and recall more instances than rank‑one or linear subspaces. The results also show that manifold steering can match the best methods when applicable, and that no method consistently outperforms prompting for these complex concepts.

arXiv Machine Learning
1d ago

Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models

The paper extends mechanistic interpretability of large language models by modeling concepts as low‑dimensional non‑linear manifolds rather than linear subspaces. It introduces a concept‑based alignment (CBA) score to compare these manifolds across layers and models, revealing block structures in intermediate layers, a shift from syntax‑dominated to mixed syntactic‑semantic concepts, and training‑dependent multilingual sharing. The study also shows that alignment patterns differ across model families and training stages, with adjacent stages aligning more closely than distant ones.

By Tido Specht, Elias Benedict Krey, Nils Neukirch, Nils Strodthoff