arXiv Computation and Language

Inter-Rater Reliability of LLM and Rule-Based Annotation for Inferential Narrative Features: Three Studies on a Turkish Corpus

arXiv Computation and Language
Aug 27

Unmatched Does Not Mean False: Incomplete Reference Sets Can Reverse Calibration Rankings in Open-Ended Theory-of-Mind Tracking

The paper demonstrates that open‑ended Theory‑of‑Mind trackers can produce valid beliefs that are absent from finite reference sets, and that treating unmatched outputs as false can reverse model‑selection rankings. By recoding references for 259 beliefs, the authors show a dramatic drop in weighted prevalence and a reversal of strictly proper Brier risk, with similar distortions observed in a 301‑question NQ‑open DPR‑BERT pipeline. The study further reveals that 90‑96% of audited unmatched beliefs are literally true, and introduces a TriSource‑Restore method that anchors reference labels to a probability‑sampled human pilot to restore calibration and ranking integrity.

By Zhexi Feng, Wuxi Chen, Bingrui Zhang
arXiv AI
Aug 26

Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes

The paper audits a 366‑day autobiographical book generated by a large language model (LLM) against an independent verification corpus. Using a four‑level rubric, 354 of the 366 days (96.7%) failed verification, with only 12 days containing corroborated scenes and 19 days containing actively contradicted claims. Regenerating the same days with current models yielded 100% verification failure, while grounding the generation in the subject’s own corpus improved the rate to 83.3% but still left substantial residual failure.

By Heather Renze
arXiv Machine Learning
Sep 14

What an odour descriptor corpus can and cannot measure: valence, attenuation, and the ceiling of the public record

The study evaluates the reliability of shared odor descriptor words across four public corpora from Pyrfume, finding substantial disagreement (I² = 80 %) and limited agreement on descriptor application (median tetrachoric = 0.795, κ = 0.413). Only a fraction of the achievable variance in odor perception is captured by current models and descriptor sets, with valence emerging as the primary missing component. Even with extensive model capacity and merged corpora, the gap remains, indicating that valence must be measured directly to improve machine olfaction.

By Stylianos Kampakis, Fabio Rovai
arXiv AI
Sep 4

CASCADE: A Component Ablation and Corpus Audit of a Layered Local Defense for MCP-Based Systems

The paper evaluates CASCADE, a fully local layered defense for Model Context Protocol (MCP)-based systems, by conducting a component ablation and corpus audit on a fixed 5,000-sample dataset. It demonstrates that the choice of aggregation convention heavily influences reported metrics, that detection performance varies with provenance, and that the released configuration does not fully disclose the operating point. The study also shows that a local review model invoked for a third of requests does not alter classification outcomes, highlighting the importance of reproducibility and transparency in defense evaluations.

By \.Ipek Abas{\i}kele\c{s} Turgut, Edip G\"um\"u\c{s}
arXiv Machine Learning
Jun 25

How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated ASR Scoring

arXiv:2606. 25487v1 Announce Type: cross Abstract: Almost every paper on LLM jailbreaks and prompt injection reports an attack-success rate (ASR), and that number is assigned not by people but by an automated judge: either a safety classifier trained for the task, or a general chat model prompted to grade.

By Yang Gao (Veyon Solutions)