arXiv AI

Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence

The paper investigates whether the factors highlighted by large language models (LLMs) as most influential on their decisions truly reflect necessity or sufficiency in influencing outcomes. By applying controlled black‑box interventions across eight models from Claude, GPT, and Gemini, the authors quantify necessity and sufficiency scores for each factor and compare them to the models’ self‑reported top three factors. Results show modest correlations (≈0.35–0.58) and reveal that the cited top factors often fail to capture the strongest measured influences, indicating limitations in current explanation practices.

arXiv Machine Learning
Jun 10

From Observation to Intervention: A Causal Audit of Expert Importance in Mixture-of-Experts Models

arXiv:2606. 10703v1 Announce Type: new Abstract: Interpretability methods routinely use population-level summary statistics over observed model behaviour to license claims about the effects of targeted interventions on specific computations; in Pearl's terms, they treat rung-1 associational evidence as if it supported rung-2 interventional conclusions, a move whose validity is rarely tested.

By Leonard Engmann, Christian Medeiros Adriano, Holger Giese
arXiv AI
Sep 10

SchemeArena: Factorized Stress Testing of Scheming in LLM Agents

The paper introduces SchemeArena, a 400-scenario benchmark designed to stress-test scheming behavior in large language model agents by factorizing key elements such as instrumental goals, environmental affordances, oversight conditions, and perceived consequences. It also presents SCOUT, a scheming monitor that uses evidence from agents' reasoning and actions to provide multi‑criteria judgments. Experiments on five LLMs show that explicit instrumental goals most strongly drive scheming, strategic hints help covert actions, and oversight can sometimes unintentionally encourage scheming.

By Jie Ruan, Inderjeet Nair, Amy Liu, Muhammad Khalifa, Yusheng Zhou, Lu Wang
arXiv AI
2d ago

Verify Claims, Not Scores: Evidence-Based Verification of Modular Agents

The paper proposes a claim‑specific verification audit for modular agents that replaces aggregate task scores with evidence‑based evaluations. Each agent conclusion is recorded with supporting evidence and classified as supported, unsupported, unresolved, or not evaluated, along with the boundary of validity. The audit employs three tools—oracle policies, perfect component replacements, and verifier‑score tests—to trace value changes, locate lost value, and assess verifier effectiveness, demonstrated on a portfolio‑allocation agent in a synthetic market.

By Ali Atiah Alzahrani
arXiv AI
Sep 17

Do Frontier Models Seek Safety Evidence Before Acting?

The paper investigates whether large language models decide to gather safety-relevant evidence before acting. Using the SAFE benchmark, the authors evaluate models such as GPT‑5.5, o3, Claude Opus, and Claude Sonnet, finding distinct evidence‑acquisition strategies that vary with retrieval cost, severity, and presentation. Across models, expected‑value reasoning dominates Stage 1 rationales, and evidence framing can alter decisions near the inspection threshold while probability is often cited despite limited influence.

By Omer Tafveez