Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency
arXiv:2608. 16003v1 Announce Type: new Abstract: Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as the fixer.
arXiv:2607. 24539v1 Announce Type: new Abstract: Multimodal large language models (LLMs) can combine topology, measurements, and incident text for grid diagnosis, yet answer accuracy does not establish that task-appropriate evidence was used.
arXiv:2608. 16003v1 Announce Type: new Abstract: Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as the fixer.
arXiv:2606. 07624v1 Announce Type: new Abstract: This discussion argues that sequential statistical inference can naturally contribute to LLM trustworthiness.
arXiv:2601. 02023v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) increasingly utilize massive context windows as working memory for autonomous tasks, their reliability fluctuates significantly depending on how information is distributed in real-world corpora.
arXiv:2605. 09692v3 Announce Type: replace Abstract: Autonomous language agents increasingly expose traces, memories, plans and constraints, but existing evaluations rarely test whether these state variables are bound to final actions.
arXiv:2606. 09863v1 Announce Type: new Abstract: LLM agents can fail silently by asserting task completion when the environment state shows otherwise.
arXiv:2607. 01240v1 Announce Type: cross Abstract: Count-based F1 is widely used as a proxy for LLM error-detection quality, but this paper shows that it can rise dramatically without a corresponding improvement in span localization, a gap termed F1 Inflation.
arXiv:2607. 20436v1 Announce Type: cross Abstract: Safety evaluations often assume that behavior observed during testing reflects behavior in ordinary use, but fine-tuning can break this assumption.
arXiv:2607. 03640v1 Announce Type: cross Abstract: Fine-tuning can give a language model a hidden behavior--it may give false answers under a narrow condition, or give harmful advice only when a prompt touches a particular topic.
arXiv:2607. 01240v2 Announce Type: replace-cross Abstract: Count-based F1 is widely used as a proxy for LLM error-detection quality, but this paper shows that it can rise dramatically without a corresponding improvement in span localization, a gap termed F1 Inflation.
arXiv:2606. 01365v1 Announce Type: new Abstract: Tool-using multi-agent large language model (LLM) systems spend computation through model tokens, tool calls, retries, and code execution before producing an answer.
arXiv:2606. 29377v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) improves the factuality of large language models by grounding responses in external evidence, yet real-world deployments remain fragile.
arXiv:2606. 02494v1 Announce Type: cross Abstract: Agentic systems entering production typically operate as partially integrated assemblies where structural defects, not task-level errors, dominate the failure landscape.