Evidence-Ledger Adjudication for Claim-Evidence Traceability
arXiv:2607. 26512v1 Announce Type: new Abstract: AI agents can draft claims faster than authors can check whether the cited or retrieved evidence supports them.
The paper introduces ARGUS, a language‑model pipeline that audits evidence for identification assumptions in difference‑in‑differences studies of climate policy. ARGUS evaluates reported evidence against an eleven‑dimension rubric, abstaining when evidence cannot be retrieved. In tests, ARGUS detects 73% of injected flaws versus 18% for a keyword approach, abstains on about 40% of assessments in 26 economics papers, and often assigns higher risk than human labels in a five‑paper pilot.
arXiv:2607. 26512v1 Announce Type: new Abstract: AI agents can draft claims faster than authors can check whether the cited or retrieved evidence supports them.
arXiv:2606. 27383v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as research assistants, yet it remains unclear whether they can calibrate research takeaways to the strength and scope of the supporting evidence.
arXiv:2607. 02586v1 Announce Type: new Abstract: Governance frameworks ask AI providers and auditors for documented evaluation evidence, and perturbation-based construct-validity audits are a common form of that evidence.
arXiv:2508.15754v2 Announce Type: replace-cross Abstract: Tool use is often assumed to monotonically improve reasoning, where external evidence is expected to help when relevant and be ignored when i...
arXiv:2609.06147v1 Announce Type: cross Abstract: Ask a language model the same question about the same document twenty times, and it sometimes returns two different answers. We built Probity, a benc...
arXiv:2608. 20116v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in settings where textual summaries, numerical observations, and external tool outputs may provide conflicting evidence.
arXiv:2606. 09809v1 Announce Type: new Abstract: AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs.
arXiv:2602. 08939v2 Announce Type: replace Abstract: Large language models increasingly produce fluent causal explanations, yet they often fail in ways aggregate accuracy cannot diagnose: confusing association with intervention, abandoning correct judgments under pressure, over-refusing valid claims, or answering when evidence is underdetermined.
The paper introduces EquiReview‑R, an AI‑assisted review system that treats omission and over‑critique as distinct risks and refines a structured concern set using evidence‑linked reasoning. It demonstrates that more criticism does not guarantee a better review, showing that many high‑recall reviews lack definitive evidence for concerns and that revision before further search is essential. On a held‑out set of papers, EquiReview‑R meets non‑inferiority for major omission, cuts major over‑critique from 15.5 % to 8.1 %, and stops on 52.4 % of papers, with gains attributed to revision rather than extra inference.
arXiv:2610.00969v1 Announce Type: cross Abstract: Large language models (LLMs) have been increasingly used for financial document analysis, including earnings call transcripts (ECTs). Beyond generati...
The study evaluates 12 instruction‑tuned open‑weight LLMs on six causal‑graph benchmarks, testing five prompting strategies and four confidence sources. Findings show that LLMs tend to over‑predict edges, misclassify indirect or reversed edges as direct, and exhibit high over‑confidence, while conventional confidence estimates are unreliable and agreement signals offer limited improvement. The results suggest LLMs should be used as externally validated soft causal priors rather than definitive causal‑structure evidence.
arXiv:2602.08889v2 Announce Type: replace Abstract: Quantitative risk assessment relies on structured expert elicitation to estimate unobservable properties. The Delphi method produces calibrated, au...