arXiv Machine Learning By Sachin Gupta

VERITYGATE: A Four-Gate Schema-Level Faithfulness Framework and Paired Benchmark for Grounded LLM Narrations over Structured Evidence

Read the original on arXiv Machine Learning →

VERITYGATE is a four‑gate framework that checks whether LLM‑generated narrations adhere to a fixed schema of declared evidence IDs, entities, numbers, and claim types, rather than verifying every fact in the prose. In experiments with GPT‑4o‑mini, Llama‑3.3‑70B, and Claude Sonnet 4.6 on 900 grounded‑ungrounded pairs, the framework identified high failure rates (up to 80.3% for mini claims) and demonstrated that a single repair pass can improve claim survival rates. The authors also provide code, data, and preliminary human studies to validate the rules and highlight gaps between schema compliance and correct prose. "whyItMatters":"The framework offers a systematic way to evaluate and improve the faithfulness of LLM explanations to structured evidence, revealing significant failure rates and the impact of repair strategies."

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 3

ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations

The paper introduces ClaimReceipt, a specification and verifier that checks whether a claim in an agent evaluation can be recomputed from retained evidence (sufficiency) and whether the evidence covers the entire experiment set (coverage). Using the CR‑2 verifier on 1,392 historical records, the authors demonstrate accurate reproduction of audit verdicts, non‑redundant field groups, and zero false positives on semantic faults. In a prospective CR‑3 run, the system correctly flags missing receipts and preserves coverage when private evidence is withheld, while adding minimal overhead to inference time and transaction size.

By Peiying Zhu, Sidi Chang
arXiv AI
3d ago

Hard-Gate Candidacy in a Deployed Validator Suite

The paper evaluates hard‑gate candidacy for validators in a deployed generative‑agent system by measuring how well each validator’s firing separates successful from failed builds. Across 13 validators and thousands of builds, only a few checks show statistically significant separation, while many fail to distinguish or never fire. The study highlights that skipped checks are recorded as passes, limiting detectable failure rates and underscoring the need for clearer evaluation records.

By Xin Xu
arXiv AI
Jun 8

Hierarchical Certified Semantic Commitment for Byzantine-Resilient LLM-Agent Collaboration

arXiv:2606. 07316v1 Announce Type: cross Abstract: Byzantine collaboration among large-language-model agents requires a finality-control primitive: given delivered stochastic, structured natural-language proposals, the protocol must decide whether the round supports a commit, what kind of commit, or a typed safe abort.

By Haoran Xu, Lei Zhang, Iadh Ounis, Xianbin Wang
arXiv AI
Aug 24

Calibrating Criterion Revision in LLM Agents: Failure Modes and a Trace-Anchored Protocol

The paper introduces a framework for evaluating how large language model agents revise their success criteria after failures, defining five non‑compensatory conditions that must be met for a criterion revision to be considered valid. Using the CMB‑0.1 protocol, the authors test twelve cross‑domain scenarios across four system configurations, finding that no model trial satisfies all five conditions and highlighting specific failure modes such as zero‑state reconstruction and inadequate intervention sensitivity. They propose a more stringent trace‑anchored CMB‑0.4 protocol to better isolate and measure criterion revision in future studies.

By Guodong Xu