arXiv:2606. 03331v1 Announce Type: cross Abstract: Consumer device repair is an important but underexplored testbed for large language models (LLMs).
By Atm Mizanur Rahman (University of Illinois Urbana-Champaign), Md Arid Hasan (University of Toronto), Syed Ishtiaque Ahmed (University of Toronto), Sharifa Sultana (University of Illinois Urbana-Champaign)
arXiv:2606. 16262v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as UX judges that inspect interfaces, diagnose usability problems, and propose repairs.
By Wenjie Wang, Yue Huang, Zipeng Ling, Han Bao, Hang hua, Xiaonan Luo, Yu Jiang, Shiyi Du, Yuexing Hao, Xiaomin Li, Yuchen Ma, Dianzhuo Wang, Yanfang Ye, Xiangliang Zhang
The paper introduces a roundtrip verification method for ensuring that large language models (LLMs) produce faithful formalizations of natural language statements. By formalizing a statement, translating it back to natural language, re-formalizing, and checking logical equivalence with a formal tool, the approach detects inconsistencies without needing ground-truth annotations. When inconsistencies are found, a diagnosis localizes the error to a specific translation step, and a scoped repair operator attempts to correct it. The framework is evaluated on the Texas Transportation Code and Texas Parks and Wildlife Code using Claude Opus and GPT-5, showing that diagnosis-guided scoped repair is most effective and that rules failing the equivalence check exhibit significantly more natural language inference drift.
By Daneshvar Amrollahi, Jerry Lopez, Clark Barrett
The paper evaluates how well open-weight large language models can repair Planning Domain Definition Language (PDDL) models using only LLMs. Experiments show that while the best LLM achieves an F1 score of 0.87—an improvement of 0.38 over a symbolic baseline—it still fails to reliably satisfy test constraints, with a mean test pass rate of only 0.82 and as low as 0.06 on the Thoughtful domain. The study concludes that current open-weight models cannot guarantee the necessary test constraint satisfaction for dependable automated model repair.
By Nader Karimi Bavandpour, Pascal Bercher
arXiv:2606. 29377v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) improves the factuality of large language models by grounding responses in external evidence, yet real-world deployments remain fragile.
By Soroush Hashemifar, Havva Alizadeh Noughabi, Fattane Zarrinkalam, Ali Dehghantanha
arXiv:2601. 05366v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly deployed as agents that invoke external tools through structured function calls.
By Zheng Luo, T Pranav Kutralingam, Ogochukwu N Okoani, Wanpeng Xu, Hua Wei, Xiyang Hu