arXiv:2606. 30963v1 Announce Type: cross Abstract: Repository-grounded automated repair is often reported as a single end-to-end capability, which hides distinct failure modes such as poor file targeting, incorrect patch synthesis, and failed iterative debugging.
By Mohammad Nour Al Awad, Sergey Ivanov
arXiv:2609.16936v1 Announce Type: cross
Abstract: Large language model (LLM)-powered coding agents have made rapid progress in automating software engineering tasks, yet repository-level issue resolu...
By Yunxiang Zhang, Haiquan Wang, JiaWei Guo, Hanyang Xia, Yan Chen, Tong Chen, Zhang Zhiwei, Junchen Ye
arXiv:2605.26380v2 Announce Type: replace-cross
Abstract: Frontier multimodal large language models (MLLMs) have been reported to achieve over 90\% accuracy on fine-grained perception benchmarks. How...
By Jingru Chen, Yiming Liu, Mingtao Chen, Sijie Chen, Richeng Xuan, Liang Yang, Zhichao Hu, Fanyang Lu
The paper introduces TED (Text-Axis Evidence Decomposition), a post‑hoc scoring method that improves anomaly localization in CLIP‑based detectors without altering the backbone or prompts. TED evaluates whether ambiguous responses are better supported by defect patches or normal patches, thereby distinguishing true defects from visually complex normal regions. Experiments show that TED significantly enhances pixel‑level localization across frozen VLM backbones and adapted hosts, especially under hard‑false‑positive competition.
By JinYoung Kim, Geonho Kim, GiJeong Park, Geonu Lee, YoungJoon Yoo
arXiv:2609.06245v1 Announce Type: cross
Abstract: Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual question answering, yet they often str...
By Yixin Wan, Tianle Zheng, Kai-Wei Chang
arXiv:2608. 07435v1 Announce Type: cross Abstract: Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to identify.
By Zixuan Lan, Luzhe Sun, Matthew R. Walter, Jiawei Zhou