arXiv AI By Zengqing Wu, Chuan Xiao

Proxy reliance in large language model decisions is uncalibrated to predictive evidence

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Machine Learning
Jul 7

Faithfulness to Refusal: A Causal Audit of Neuron Selectors

arXiv:2607. 05355v1 Announce Type: cross Abstract: Attribution scores increasingly identify which neuron rows of a language model matter for applications such as pruning, interpretability, and editing for safety, yet whether they identify causally important rows is rarely tested directly.

By Ananth Eswar, Pratinav Seth, Utsav Avaiya, Vinay Kumar Sankarapu
arXiv AI
2d ago

Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias

The paper investigates whether language models still encode occupational biases even when they appear unbiased in behavioral tests. Using a causal framework, the authors separate bias into internal representations of user competence and observable outputs, deriving steering vectors that show these representations influence model behavior in question‑answering and hiring tasks. Across several open‑weight models, demographic factors such as gender, race, and socioeconomic status affect the models’ internal competence representations, revealing hidden bias that behavioral metrics alone may miss.

By Keren Fuentes, Aaron Mueller