arXiv:2607. 14707v1 Announce Type: cross Abstract: Large language models routinely produce fluent answers to single-shot prompts, yet deploying them as reliable components of a domain decision system is substantially harder.
By Akash Raj
arXiv:2607. 03656v1 Announce Type: cross Abstract: Large Language Models are increasingly used to turn natural-language requirements into code.
By Adarsh Vatsa, Sachi Shome, Yingming Zhou, William Eiers
arXiv:2607. 28229v1 Announce Type: cross Abstract: The web is increasingly accessed by AI agents rather than humans.
By Luigi Sigillo, Matteo Silvestri, Francesco Tabaro, Rajat Bhatnagar, Syed Irtaza Mubashar, Matt Jeffryes, Daljit Nijjer, Vittorio Perera, Ola Spjuth, Julio Saez-Rodriguez, Melissa Harrison, Fabio Petroni
arXiv:2607. 01256v1 Announce Type: cross Abstract: Overwhelmed courts in the United States review millions of default judgments each year.
By Theodora Worledge, Othman Bensouda Koraichi, Daniel Bernal, Aviv Caspi, Tatsunori Hashimoto, Carlos Guestrin, David Freeman Engstrom
The paper introduces Tasks over Application Manuals (TAM), a benchmark designed to test long‑horizon procedural reasoning in large language models. TAM uses real‑world tasks from ICD‑10‑CM clinical coding and U.S. federal sentencing, requiring models to follow extensive, rule‑based manuals and perform interdependent steps to produce exact answers. Experiments with GPT‑5 and various prompting strategies show very low exact‑match accuracy—1% for coding and 15.5% for sentencing—highlighting a gap between current benchmarks and the ability to reliably follow complex procedures.
By Utkarsh Soni, Syed Shariyar Murtaza, Yifan Nie, Sachin Chandrasekhar, Eugene Wen
The paper introduces TRACE, a fine‑tuning framework for Retrieval‑Augmented Generation (RAG) that addresses conflicts between retrieved knowledge and a model’s internal knowledge. TRACE uses multi‑agent debate traces to identify correct and incorrect candidates and answer‑shift patterns, providing fine‑grained supervision for reliable knowledge‑source selection. It also incorporates an answer‑completeness regularization mechanism to prevent empty, overly short, or prematurely terminated responses, thereby improving robustness against misleading retrieved content and enhancing answer quality.
By Zhengchen Huang, Yundong Sun, Minrui Song, Shuanglong Yao, Ye Liu, Ji Chen, Xing Wang