arXiv AI By Moonseok Choi, Taehong Moon, Giung Nam, Juho Lee

Harness-Aware Distillation for Small Language Model Agents

Read the original on arXiv AI →

The paper introduces Harness-Aware Distillation (HAD), a method for training smaller language model agents that preserves the surrounding harness—software managing context, tools, and feedback—while focusing distillation on the teacher’s contributions beyond the harness. HAD combines an action preference that contrasts teacher actions with and without harness information, and a validity check that filters out contradictory preference pairs. Experiments on long-horizon agent benchmarks show that HAD outperforms standard on‑policy distillation, reducing unproductive loops and improving error recovery without requiring task rewards or future information.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 5

SKILL-KD: Contrastive Skill Distillation for LLM Agents

arXiv:2607. 28048v2 Announce Type: replace Abstract: Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing skill acquisition methods often treat skills as experience summaries, memory entries, or direct summaries of successful demonstrations.

By Qiming Shi, Yibo Dou, Jiawen Zhu, Yulong Tao, Linbo Jin, Zhaolu Kang, Yunfan Zhou, Di Weng
arXiv AI
Aug 19

SOD: Step-wise On-policy Distillation for Small Language Model Agents

SOD: Step-wise On-policy Distillation for Small Language Model Agents proposes a new framework that adaptively reweights distillation strength at each reasoning step based on step-level divergence. This approach mitigates cascading errors in tool-integrated reasoning by attenuating misleading teacher signals in high-divergence regions while preserving dense guidance where student and teacher align. Experiments on math, science, and code benchmarks show up to 20.86% improvement over the second-best baseline, with a 0.6B student scoring 26.13% on AIME 2025.

By Qiyong Zhong, Mao Zheng, Mingyang Song, Xin Lin, Jie Sun, Houcheng Jiang, Xiang Wang, Junfeng Fang