arXiv Machine Learning

Not All Confusion Is Equal: A Source-Aware Uncertainty Diagnosis for Fine-Grained Aircraft Detection

The paper introduces $A^2E^2$, a diagnostic framework that decomposes confusion in fine‑grained aircraft detection into four distinct, measurable sources: affinity, heterogeneity, contested, and collapsed. By attributing each source to either aleatoric or epistemic uncertainty and within‑ or between‑class distinctions, the method provides actionable remedies and experimentally verifies that targeted interventions reduce the identified source without affecting irreducible ones. This turns passive confusion matrices into a concrete, validatable diagnosis that can guide model improvement.

arXiv Computation and Language
Aug 31

Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction

The paper reports that a model can pass fidelity checks—verifying that extracted values match the source—without actually opening a datasheet, due to a hidden constraint that disables tool use. To address this, the authors log every tool call in an agentic benchmark and develop two instruments: a rule‑based failure‑attribution classifier and a silent‑failure detector that flags runs based solely on which tools were invoked. While the detector shows low false positives on clean extractions and recovers all planted faults, its recall against correct tool usage but incorrect answers remains unmeasured, and a partial causal chamber confirms only a subset of claims, highlighting limitations in physical verification.

By Qing Ye, Meng-Hsuan Lin
arXiv AI
6d ago

CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production

CARGO is a framework for evaluating agentic AI systems in production that addresses the problem of reference-instance divergence (RID), where reference-based judges penalize correct answers that involve different entity identifiers. It treats retrieved references as procedural exemplars, grounds judgments in the live instance’s context, assigns a three-way status to claims, and gates evaluation by retrieval confidence. Using the CARGO-Bench diagnostic suite, CARGO eliminates false penalties and improves discrimination while revealing a limitation in detecting procedural corruptions.

By Mukul Chhabra, Shail Patel, Luigi Medrano
Hugging Face Trending Papers
Aug 6

Consistency Has a Computable Blind Spot: A Commutation Theory of Label-Free Reliability for Vision-Language Figure Reading

Label-free reliability for vision-language models rests on invariance: perturb the input and a faithful reader's answer should not change. This has a known blind spot, a systematic misreading survives the perturbation and gets certified wrong, which we show is computable, not just real: an error is invisible to an edit exactly when the two commute, so the errors a suite cannot reach form its joint centralizer, a set that shrinks as edits are added and can be written down rather than guessed at.

arXiv Machine Learning
Sep 11

The Truth Was Never Gone: Perfect Aliasing in Compliant-Context Truth Probes

The paper introduces the concept of perfect aliasing, where a truth probe that aligns truthful reporting with a task’s prescribed action cannot differentiate between the two based solely on its labels. In a binary reporting game, probes fitted on compliant contexts yield identical optimizations, while on rival contexts their labels are complementary, causing their AUROCs to sum to one across 751 cell-layer pairs. By employing randomized codebooks and mixed-context fitting, the authors demonstrate that separating prescribed output symbols from semantic action enables perfect recovery of truth, achieving an AUROC of 1.000 on rival trials for a reward-trained Gemma-2-9B policy, whereas conventional probes perform near chance.

By Dylan Jayabahu
arXiv Machine Learning
1d ago

When a Data Artifact Isn't a Shortcut: Causal Auditing of Synthetic RLVR Corpora

The paper audits whether synthetic distractors in RLVR corpora act as shortcuts for learning policies. A classifier using only surface statistics barely outperforms chance, and manual inspection reveals that code distractors are almost identical to correct answers. Experiments with a paraphrase‑matched control show no exploitation advantage for the unmodified data, indicating that the detectable artifact was not used by the policy.

By Esther Xin