arXiv:2607. 08065v1 Announce Type: new Abstract: LLM-as-judge (Zheng et al.
By Kaihua Ding
arXiv:2607. 23386v1 Announce Type: new Abstract: We document a failure class in frontier large language models -- exception chain collapse -- observed in eligibility evaluation under nested conditional rules of the form "A is required UNLESS B applies, UNLESS C overrides B".
By Paul Simpson, John Kozak, Lisa Doake
arXiv:2609.17545v1 Announce Type: new
Abstract: Deep learning models for cervical cytology are almost always evaluated as if every prediction must be acted upon, yet a screening system deployed along...
By Nisreen Albzour, Sarah S. Lam
The paper introduces TrustSwap, a counterfactual test that swaps or removes source reliability labels while keeping evidence text constant, to evaluate how retrieval‑augmented fact‑checking models respond across verdict, confidence, and search decisions. Experiments on untrained and RL‑trained models show that confidence and search largely follow labels, yet label changes can flip a significant portion of verdicts, especially in larger models. The authors propose trust‑swap augmentation (TSA) to mitigate this shortcut, demonstrating reduced verdict flip rates and maintained accuracy in several settings, though its effectiveness diminishes at larger model scales.
By Jianchang Su, Yiwei Yang, Wei Zhang
arXiv:2606. 29484v1 Announce Type: cross Abstract: Modern deepfake detectors are rarely consumed as bare classifiers.
By Md Anas Biswas
arXiv:2608. 12652v1 Announce Type: cross Abstract: Benchmark contamination is diagnosed today with n-gram overlap, with likelihood-based membership inference, or with canary strings, and each needs something usually unavailable: the training corpus, a well-chosen test statistic, or foresight at dataset release.
By Florian Braun