EviScope is a new paired counterfactual benchmark that evaluates grounded language models by fixing the question while manipulating evidence—adding, removing, distracting, or contradicting it. The v1.1 dataset includes 40 four‑condition quartets with repaired counterfactual claims and span‑level support labels for automated assessment. Experiments on Qwen2.5‑7B, Llama 3.1 8B, and Gemini 3.5 Flash show that paired metrics reveal grounding behaviors hidden by simple answer accuracy, such as unsupported answers, conflict blindness, and incorrect non‑answer actions.
By Suryadeep Singh Deswal
The paper investigates how memory systems can answer a current query correctly yet fail to retain distinctions needed for later updates. Using a paired‑history audit, the authors evaluate 24 history pairs across six synthetic mechanisms and two model backends, achieving perfect reveal accuracy on DeepSeek and high accuracy on GLM. Record‑level audits reveal specific failures in structured reveal memories and frontier late‑reference adequacy, and the authors test a label‑equivariant repair that only partially restores correctness.
By Guangzhe Zhang
arXiv:2606. 27383v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as research assistants, yet it remains unclear whether they can calibrate research takeaways to the strength and scope of the supporting evidence.
By Yu Fu, Yongqi Kang, Yong Zhao
Grounded language-model systems are often evaluated by final answer accuracy, yet a correct answer can be unsupported, drawn from the wrong source, or produced when evidence is insufficient or contrad...
arXiv:2608.29582v1 Announce Type: cross
Abstract: Current evaluations of large language models (LLMs) primarily focus on factual knowledge retrieval, overlooking the fundamental challenge of navigati...
By Yi Yu, Bo Wang, Chong Feng, Ge Shi, Xia Liu, Ziyi Yang, Xuewen Shi
The paper introduces a joint fact‑verification score that evaluates both answers and the evidence submitted with them. On the FEVEROUS dataset, replacing the DCUF evidence with UnifEE evidence improves the strict score by about 9.6 percentage points, while answer accuracy rises only 1.96 points. The study also shows that increasing context length for large language models yields modest evidence‑gain improvements, and that detailed answer‑evidence analyses uncover patterns missed by aggregate metrics.
By Han Chen, Yingrui Li
arXiv:2605. 09692v3 Announce Type: replace Abstract: Autonomous language agents increasingly expose traces, memories, plans and constraints, but existing evaluations rarely test whether these state variables are bound to final actions.
By Xiao Jia
The paper investigates how personalized agents decide to use, ignore, update, or query retrieved user memory before acting on a task. An empirical audit protocol is developed to test structured intermediate outputs, revealing that while exposing state definitions improves accuracy, an explicit state-output field does not significantly enhance policy accuracy for large language models. The study also shows that example-level accuracy overstates consistency, with full four‑way family success being rare, and that providing benchmark‑associated state labels merely conditions predictions rather than proving internal fidelity.
By Yihang Chen, Pin Qian, Su Wang, Chong Peng, Huan Xu, Shuaiting Li, Yiqi Sun
The study examines how two small instruction‑tuned language models, Qwen2.5‑1.5B and Llama‑3.2‑1B, respond to user pushback on TriviaQA. When initially correct, the models flip to a wrong answer in about 42–43% of cases, with the effectiveness of different pushback styles varying by model. Attempts to decode capitulation from the pre‑response residual stream fail under a rigorous validation protocol, revealing overfitting and a measurement hazard that underestimates capitulation by 18–24 percentage points.
By Saad Aamir, Muhammad Awais Bin Adil
arXiv:2608. 03674v1 Announce Type: new Abstract: Causal diagnostic models must explain how conclusions follow from evidence because diagnoses guide repairs and treatments.
By Jian Zhang, Bingyi Wang, Yizhi Liu
arXiv:2605. 03534v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) grounds answers in retrieved passages, yet relevance does not guarantee sufficiency: a topical passage may still fail to justify the answer.
By Jingxi Qiu, Zeyu Han, Cheng Huang
arXiv:2605. 14473v4 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) is usually evaluated by whether the final answer is correct.
By Yihang Chen, Pin Qian, Su Wang, Sipeng Zhang, Huan Xu, Shuhuai Lin, Xinpeng Wei