arXiv AI

Compiling Learning Problems into Adaptation Programs for Language Models

arXiv AI
Sep 15

Drift-Constrained Optimization: Only Direction Matters in Fine-Tuning Instruct Models

The paper introduces Drift-Constrained Optimization (DCO), a framework that treats behavioral drift during fine‑tuning of instruction models as a bounded constraint rather than an uncontrolled side effect. By defining a drift budget, the authors reformulate fine‑tuning as a direction‑selection problem, showing that choosing different update directions can qualitatively change outcomes. Experiments on Qwen3 models demonstrate that carefully selected directions improve scientific reasoning and multilingual translation while preserving reasoning capabilities and general performance.

By Fei Yuan, Changjiang Gao, Yilei Tu, Yifeng Liu, Shujian Huang, Yu Qiao
arXiv AI
Aug 24

UpgradeBench: A Decision-Centric Benchmark for Upgrading Fine-Tuned LLM Specialists

UpgradeBench is a decision‑centric longitudinal benchmark that evaluates how fine‑tuned language‑model specialists should be handled when new base‑model releases occur. It covers four consecutive Qwen releases, a continuation checkpoint, six tasks, two model sizes, and OLMo checkpoints with known training lineage, and examines whether retraining, adapter transfer, or other recovery strategies improve specialist performance. The benchmark reveals that upgrade gains vary by task and release interval, that direct adapter copying is sensitive to pretraining distance, and that teacher relabeling can recover specialists without new annotations. "whyItMatters":"The study provides actionable insights into the cost‑effective management of specialist models across model releases, showing how to balance retraining effort with performance gains."

By Ye Chen, Weining Zhang
arXiv AI
Sep 2

Predicting Program Exit Code with LLMs and Programming Language Semantics

The paper introduces Program Executability Prediction (PrEx), a task that asks large language models (LLMs) to determine whether a program is semantically valid or invalid and, if invalid, to identify the violated formal rule. To evaluate this, the authors create a dataset of systematically generated invalid programs derived from valid ones and test open‑source coding LLMs across different semantic formalisms, semantic shifts, and program splits (human‑written, LLM‑translated, fuzzer‑generated). Results show that LLMs rely more on pre‑training priors than on the provided semantics, performing poorly on modified semantics and with increasing program complexity.

By Lara Marinov, Aditya Thimmaiah, Jayanth Srinivasa, Junyi Jessy Li, Milos Gligoric
arXiv AI
Aug 11

Janus: An Algorithm-Evaluator Co-Evolution Framework for LLM-Driven Discovery under Expensive Evaluation Budgets

arXiv:2608. 08189v1 Announce Type: new Abstract: LLM-driven program discovery relies on rapid evaluator feedback, but many scientific and engineering tasks require high-fidelity simulations, hardware execution, or physical experiments, making each evaluation expensive.

By Ximeng Liu, Qianlong Wang, Yingming Mao, Annan Li, Yatao Li, Shizhen Zhao, Jianmin Wu, Dawei Yin, Dou Shen
arXiv AI
4d ago

More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses

The paper introduces a controlled evaluation to disentangle answer coverage, repeatable task advantages, and gains from pre‑execution selection in large‑language‑model (LLM) harnesses. On 386 MATH‑500 tasks, eight generated harnesses and a baseline with nine identical copies were compared over three executions each, revealing that identical programs provide a 2.16‑point repeat‑averaged oracle headroom while generated programs show more repeatable score patterns but mainly expose persistent weaknesses. The study concludes that coverage and repeatability alone cannot justify claims of useful specialization and proposes an evaluation standard for harness diversity that requires task advantages to persist across executions and improve on additional fixed‑program executions under matched inference budgets.

By Ziyang Xu, Haitian Zhong, Hao Zhou, Hao Qin, Chenhan Jin, Te Qi, Shengze Xu, Tieyong Zeng
arXiv Machine Learning
Jul 10

TTHE: Test-Time Harness Evolution

arXiv:2607. 08124v1 Announce Type: cross Abstract: The behavior of an LLM agent is determined not only by the underlying model, but also by its harness: the executable program that constructs context, invokes tools, verifies intermediate results, and recovers from failures.

By Jun Nie, Yonggang Zhang, Jun Song, Qianshu Cai, Dahai Yu, Yike Guo, Xinmei Tian, Bo Han