arXiv AI

Leveraging Metamemory Agent for Enhanced Data-Free Code Generation in Large Language Models

arXiv:2501. 07892v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown strong performance in automated code generation, with few-shot prompting widely used for its simplicity and effectiveness.

arXiv AI
Sep 2

WiseSpec: Requirements-Driven Agents for Code Generation

WiseSpec is a requirements‑driven agent framework designed to improve repository‑level code generation. It automatically builds structured, information‑rich requirements, evaluates their quality via execution‑based tests, and iteratively refines them to better guide code generation. Experiments show WiseSpec outperforms all baselines, achieving an average 13.17% improvement in %Resolved.

By Zhao Tian
arXiv AI
Jul 28

KG2Code: Bridging Knowledge Graphs and Large Language Models via Executable Code for Question Answering

arXiv:2607. 22652v1 Announce Type: new Abstract: Recent research has explored the integration of knowledge graphs (KGs) with large language models (LLMs) to enhance their performance on downstream knowledge-intensive tasks, particularly knowledge graph question answering (KGQA).

By Yike Wu, Nan Hu, Guilin Qi, Guohui Xiao, Chen Jiang, Xinchun Zou, Yuchen Lu, Songlin Zhai, Yongrui Chen, Yuyang Zhang, Xiaoguang Li, Lifeng Shang, Jiaoyan Chen, Jeff Z. Pan
arXiv Machine Learning
Jun 11

FronTalk: Benchmarking Front-End Development as Conversational Code Generation with Multi-Modal Feedback

arXiv:2601. 04203v2 Announce Type: replace-cross Abstract: We present FronTalk, a benchmark for front-end code generation that pioneers the study of a unique interaction dynamic: conversational code generation with multi-modal feedback.

By Xueqing Wu, Zihan Xue, Da Yin, Shuyan Zhou, Kai-Wei Chang, Nanyun Peng, Yeming Wen
arXiv Computation and Language
Aug 28

SPT: Skills as Pre-Training Data for Agentic Language Models

The paper introduces Skill Pre-Training (SPT), a mid‑training approach that uses public skill packages as data for agentic language models. By applying causal language modeling to a collection called SkillCorpus and employing a Reference Insert strategy to keep file relationships intact, SPT improves agentic performance across various model sizes while largely preserving general capabilities. Experiments also show that mixing skill data with general corpora yields further benefits.

By Yufei Sun, Yudong Li, Yiming Cheng
arXiv Machine Learning
Aug 31

Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks

The paper introduces CodeRQ-Bench, the first benchmark for assessing large language model reasoning quality across coding tasks such as generation, summarization, and classification. It analyzes over a thousand mismatches from existing evaluators, identifies recurring limitations, and derives design insights that lead to a new two‑stage evaluator, VERA. Experiments show VERA outperforms strong baselines, improving AUCROC by up to 0.26 and AUPRC by up to 0.21 on four datasets.

By Yuangang Li, Justin Tian Jin Chen, Ethan Yu, David Hong, Iftekhar Ahmed
arXiv AI
Jun 4

Can Generalist Agents Automate Data Curation?

arXiv:2606. 04261v1 Announce Type: new Abstract: Curating training data is among the most consequential yet labor-intensive parts of modern AI development: practitioners iteratively propose, implement, evaluate, and revise data policies against noisy benchmark feedback.

By Feiyang Kang, Hanze Li, Adam Nguyen, Mahavir Dabas, Jiaqi W. Ma, Frederic Sala, Dawn Song, Ruoxi Jia