arXiv AI By Xingyang Yu

DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory

Read the original on arXiv AI →

arXiv:2607. 23614v1 Announce Type: cross Abstract: We present DualityCert, a symbolic verifier for candidate Seiberg-duality claims in four-dimensional N=1 quiver gauge theories.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
4d ago

Multi-Channel Mitigation of Source-Trust Shortcuts in Fact-Checking RL Agents

The paper introduces TrustSwap, a counterfactual test that swaps or removes source reliability labels while keeping evidence text constant, to evaluate how retrieval‑augmented fact‑checking models respond across verdict, confidence, and search decisions. Experiments on untrained and RL‑trained models show that confidence and search largely follow labels, yet label changes can flip a significant portion of verdicts, especially in larger models. The authors propose trust‑swap augmentation (TSA) to mitigate this shortcut, demonstrating reduced verdict flip rates and maintained accuracy in several settings, though its effectiveness diminishes at larger model scales.

By Jianchang Su, Yiwei Yang, Wei Zhang
arXiv Computation and Language
Aug 31

Why Didn't It Check? Unsupported Final Claims and Their Repair in Two Tool-Equipped Language Models

The study investigates how language models equipped with tools can still produce unsupported final claims, even when a single tool call could resolve the uncertainty. It defines two metrics—occurrence (how often unsupported claims arise) and conditional repair (how often they are fixed when evidence is provided). Experiments on Qwen3-32B and Gemma 4 show that providing the missing evidence consistently repairs all unsupported claims in the Qwen3-32B setup, while the Gemma 4 model never produced unsupported claims under the tested conditions.

By Justin Bronder
arXiv Machine Learning
1d ago

VERITYGATE: A Four-Gate Schema-Level Faithfulness Framework and Paired Benchmark for Grounded LLM Narrations over Structured Evidence

VERITYGATE is a four‑gate framework that checks whether LLM‑generated narrations adhere to a fixed schema of declared evidence IDs, entities, numbers, and claim types, rather than verifying every fact in the prose. In experiments with GPT‑4o‑mini, Llama‑3.3‑70B, and Claude Sonnet 4.6 on 900 grounded‑ungrounded pairs, the framework identified high failure rates (up to 80.3% for mini claims) and demonstrated that a single repair pass can improve claim survival rates. The authors also provide code, data, and preliminary human studies to validate the rules and highlight gaps between schema compliance and correct prose. "whyItMatters":"The framework offers a systematic way to evaluate and improve the faithfulness of LLM explanations to structured evidence, revealing significant failure rates and the impact of repair strategies."

By Sachin Gupta