arXiv AI

From Execution to Capability: Scientific Experience Consolidation via Procedural Knowledge Synthesis

arXiv:2607. 24459v1 Announce Type: new Abstract: Large language models increasingly solve scientific-computing tasks, but executable feedback from one problem rarely becomes durable capability on subsequent problems.

arXiv AI
Sep 3

SkillGLoW: Procedural-Family Skill Consolidation for Self-Improving Agents on Long-Horizon Task Streams

SkillGLoW introduces a new way for large language model agents to self‑improve by consolidating procedural skills shared across related tasks. Instead of storing all skills in a single global document or a flat per‑task pool, SkillGLoW aggregates local skills into procedural families, compresses them into de‑instantiated global priors, and regenerates instance‑specific details on demand. Experiments on four diverse benchmarks show that these priors improve performance by an average of 17.2 points over a no‑skill baseline, are more compact than per‑task pools, and enable better transfer to unseen tasks.

By Ao Yan, Xin Zhang, Jiawei Du, Joey Tianyi Zhou
arXiv AI
Jun 26

NebulaExp-8B: An Empirical Post-Training Pipeline via Full-Scale Ablation Research

arXiv:2606. 26671v1 Announce Type: new Abstract: Post-training alignment determines the reasoning and human preference following capabilities of large language models, yet most existing works withhold detailed data construction, filtering rules and training recipes, which hinders community reproducibility and lightweight model optimization.

By Qiaobo Hao, Yangqian Wu, Shunyi Wang, Zhongjian Zhang, Ziqun Li, Yayin He, Muqing Li, Chen Zhong
arXiv AI
Sep 18

SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science

SCICONVBENCH is a benchmark designed to evaluate large language models (LLMs) on multi‑turn clarification tasks in computational science. It focuses on two key abilities: eliciting missing information (disambiguation) and resolving contradictory requests (inconsistency resolution) across four domains—fluid mechanics, solid mechanics, materials science, and partial differential equations. The benchmark pairs a structured task ontology with a rubric‑based evaluation framework, measuring LLM performance in clarification behavior, conversational grounding, and final‑specification fidelity, and reveals that even top models only resolve about 52.7% of disambiguation cases in fluid mechanics while often making ungrounded assumptions.

By Nithin Somasekharan, Youssef Hassan, Shiyao Lin, Gihan Panapitiya, Patrick Emami, Anurag Acharya, Sameera Horawalavithana, Shaowu Pan
arXiv Machine Learning
Jun 2

ATLAS: Agentic Test-time Learning-to-Allocate Scaling

arXiv:2606. 01667v1 Announce Type: new Abstract: Test-time scaling has become a major way to improve large language model reasoning, but its orchestration has remained designer-engineered: a fixed sample budget, a fixed refinement loop, a fixed scoring rule, or a fixed search policy decides how compute is spent, leaving the model in charge of solving but not of orchestration.

By Peijia Qin, Qi Cao, Pengtao Xie
arXiv AI
Jun 30

Hierarchical Experimentalist Agents

arXiv:2606. 29315v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to take actions in the real world and support human decision-making, yet most agents rely on parametric knowledge, fixed post-training data, retrieval, or search.

By Abhranil Chandra, Sankaran Vaidyanathan, Utsav Dhanuka, Varun Gandhi, Scott Niekum
arXiv AI
Sep 25

SciWalker: Synthesizing Scientific Coding Problems with Operator Graphs and Execution Feedback

SciWalker is a framework that automatically synthesizes scientific coding problems by sampling operator chains from scientific library interfaces and using execution feedback to refine generated problem statements, solutions, and tests. It produces 8,178 high‑quality problems across five scientific domains and 32 subdomains, and training a large language model with these problems improves its scientific coding accuracy by nearly 10 percentage points. The approach combines structured workflow composition with verification and quality review to enable scalable, scientifically grounded task generation.

By Chenxi Li, Wenxuan Zeng, Yun Luo, Fangchen Yu, Peng Ye, Yu Cheng, Jun Zhang
arXiv AI
3d ago

RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models

RankEvolve is an auto‑research framework that evolves generative ranking models by orchestrating multiple large‑language‑model coding agents through an Executable Operating Protocol (EOP). The system compiles a state machine that enforces phases, gates, branches, and loops, while a meta‑meta‑harness lets agents review and repair each other’s code. In budget‑matched experiments, heterogeneous composition of agents raised execution accuracy from 45.8 % to 62.5 % and reduced silent critical‑defect rates, achieving notable gains on the HSTU recommender and other benchmarks.

By Zheng Chen, Linfeng Liu, Hong Li, Hong Yan