arXiv AI By Liwei Dong, Jiahao Zhao, Nan Xu

From Execution to Capability: Scientific Experience Consolidation via Procedural Knowledge Synthesis

Read the original on arXiv AI →

arXiv:2607. 24459v1 Announce Type: new Abstract: Large language models increasingly solve scientific-computing tasks, but executable feedback from one problem rarely becomes durable capability on subsequent problems.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 3

SkillGLoW: Procedural-Family Skill Consolidation for Self-Improving Agents on Long-Horizon Task Streams

SkillGLoW introduces a new way for large language model agents to self‑improve by consolidating procedural skills shared across related tasks. Instead of storing all skills in a single global document or a flat per‑task pool, SkillGLoW aggregates local skills into procedural families, compresses them into de‑instantiated global priors, and regenerates instance‑specific details on demand. Experiments on four diverse benchmarks show that these priors improve performance by an average of 17.2 points over a no‑skill baseline, are more compact than per‑task pools, and enable better transfer to unseen tasks.

By Ao Yan, Xin Zhang, Jiawei Du, Joey Tianyi Zhou
arXiv AI
Jun 26

NebulaExp-8B: An Empirical Post-Training Pipeline via Full-Scale Ablation Research

arXiv:2606. 26671v1 Announce Type: new Abstract: Post-training alignment determines the reasoning and human preference following capabilities of large language models, yet most existing works withhold detailed data construction, filtering rules and training recipes, which hinders community reproducibility and lightweight model optimization.

By Qiaobo Hao, Yangqian Wu, Shunyi Wang, Zhongjian Zhang, Ziqun Li, Yayin He, Muqing Li, Chen Zhong
arXiv AI
Sep 18

SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science

SCICONVBENCH is a benchmark designed to evaluate large language models (LLMs) on multi‑turn clarification tasks in computational science. It focuses on two key abilities: eliciting missing information (disambiguation) and resolving contradictory requests (inconsistency resolution) across four domains—fluid mechanics, solid mechanics, materials science, and partial differential equations. The benchmark pairs a structured task ontology with a rubric‑based evaluation framework, measuring LLM performance in clarification behavior, conversational grounding, and final‑specification fidelity, and reveals that even top models only resolve about 52.7% of disambiguation cases in fluid mechanics while often making ungrounded assumptions.

By Nithin Somasekharan, Youssef Hassan, Shiyao Lin, Gihan Panapitiya, Patrick Emami, Anurag Acharya, Sameera Horawalavithana, Shaowu Pan