arXiv:2608. 16003v1 Announce Type: new Abstract: Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as the fixer.
By Parsa Mazaheri, Kasra Mazaheri
EviScope is a new paired counterfactual benchmark that evaluates grounded language models by fixing the question while manipulating evidence—adding, removing, distracting, or contradicting it. The v1.1 dataset includes 40 four‑condition quartets with repaired counterfactual claims and span‑level support labels for automated assessment. Experiments on Qwen2.5‑7B, Llama 3.1 8B, and Gemini 3.5 Flash show that paired metrics reveal grounding behaviors hidden by simple answer accuracy, such as unsupported answers, conflict blindness, and incorrect non‑answer actions.
By Suryadeep Singh Deswal
arXiv:2606. 07624v1 Announce Type: new Abstract: This discussion argues that sequential statistical inference can naturally contribute to LLM trustworthiness.
By Yao Xie
InSituMeasure is a new benchmark that tests multimodal large language models on situated measurement tasks in industrial scenes. It includes 2,922 real monitoring images of eight types of engineering instruments, with detailed gauge-attribute labels and noise tags for failure diagnosis. The benchmark defines metrics for numerical accuracy, unit consistency, rejection of unanswerable queries, and alignment of model failures with annotated error factors, revealing that even the best models achieve only 25.7% joint value‑unit accuracy and 51.8% confidence‑diagnosis F1.
By Chao Shen, Xinyuan Li, Yunfan Zhou, Jianguo Yao, Haibing Guan, Zhihai Wang, Xijun Li
The paper audits a developer‑accessible on‑device language model, revealing that it can confidently produce incorrect answers while refusing benign prompts, a phenomenon termed task‑asymmetric miscalibration. The model’s confident outputs are surface‑indistinguishable, with classifiers based on user‑visible features failing to separate correct from wrong responses. The authors propose a model‑agnostic audit protocol, a surface‑indistinguishability test, and a black‑box consistency wrapper that improves reliability without requiring model access.
By Shashwat Pandey, Satwik Pandey, Suresh Raghu
arXiv:2609.38465v1 Announce Type: cross
Abstract: Unified multimodal models (UMMs) are increasingly designed around gradient conflict between understanding and generation objectives. The premise that...
By Shuyang Jiang, Fucheng Deng, Yuchuan Luo, Zhenyu Wu