arXiv Machine Learning

Same Evidence, Different Target: Decoding How Diagnostic Evidence Bears on Causal Questions from Language-Model States

arXiv:2607. 26929v1 Announce Type: cross Abstract: The same diagnostic result can support or challenge one causal claim yet fail to address another when the claims concern different populations, outcomes, estimands, pathways, or identifying assumptions.

arXiv Computation and Language
Sep 16

EviScope: Paired Counterfactual Evidence Diagnostics for Faithful and Efficient Grounded Language Models

EviScope is a new paired counterfactual benchmark that evaluates grounded language models by fixing the question while manipulating evidence—adding, removing, distracting, or contradicting it. The v1.1 dataset includes 40 four‑condition quartets with repaired counterfactual claims and span‑level support labels for automated assessment. Experiments on Qwen2.5‑7B, Llama 3.1 8B, and Gemini 3.5 Flash show that paired metrics reveal grounding behaviors hidden by simple answer accuracy, such as unsupported answers, conflict blindness, and incorrect non‑answer actions.

By Suryadeep Singh Deswal
arXiv AI
Sep 18

Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression

The paper investigates how memory systems can answer a current query correctly yet fail to retain distinctions needed for later updates. Using a paired‑history audit, the authors evaluate 24 history pairs across six synthetic mechanisms and two model backends, achieving perfect reveal accuracy on DeepSeek and high accuracy on GLM. Record‑level audits reveal specific failures in structured reveal memories and frontier late‑reference adequacy, and the authors test a label‑equivariant repair that only partially restores correctness.

By Guangzhe Zhang
arXiv Computation and Language
Sep 24

What Changes When Fact-Verification Scores Improve? Evidence and Answer Accounting Across Trained Verifiers and LLMs

The paper introduces a joint fact‑verification score that evaluates both answers and the evidence submitted with them. On the FEVEROUS dataset, replacing the DCUF evidence with UnifEE evidence improves the strict score by about 9.6 percentage points, while answer accuracy rises only 1.96 points. The study also shows that increasing context length for large language models yields modest evidence‑gain improvements, and that detailed answer‑evidence analyses uncover patterns missed by aggregate metrics.

By Han Chen, Yingrui Li
arXiv AI
Aug 19

Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification

The paper investigates how personalized agents decide to use, ignore, update, or query retrieved user memory before acting on a task. An empirical audit protocol is developed to test structured intermediate outputs, revealing that while exposing state definitions improves accuracy, an explicit state-output field does not significantly enhance policy accuracy for large language models. The study also shows that example-level accuracy overstates consistency, with full four‑way family success being rare, and that providing benchmark‑associated state labels merely conditions predictions rather than proving internal fidelity.

By Yihang Chen, Pin Qian, Su Wang, Chong Peng, Huan Xu, Shuaiting Li, Yiqi Sun
arXiv Machine Learning
Sep 17

No Usable Linear "Capitulation Direction" in Two Small LLMs: A Validation Protocol for Activation-Steering Claims, and a Cross-Family Behavioral Study of Sycophancy Under Pushback

The study examines how two small instruction‑tuned language models, Qwen2.5‑1.5B and Llama‑3.2‑1B, respond to user pushback on TriviaQA. When initially correct, the models flip to a wrong answer in about 42–43% of cases, with the effectiveness of different pushback styles varying by model. Attempts to decode capitulation from the pre‑response residual stream fail under a rigorous validation protocol, revealing overfitting and a measurement hazard that underestimates capitulation by 18–24 percentage points.

By Saad Aamir, Muhammad Awais Bin Adil