arXiv AI

STEMMA: An Adversarial Multi-Agent Framework for Evaluating Self-Identity Consistency in LLMs

arXiv:2608. 08164v1 Announce Type: cross Abstract: Knowledge Distillation is a widely adopted technique in the training and fine-tuning of large language models (LLMs) enabling transfer of structured information and functional behavior from a large teacher model to a smaller student model while significantly reducing computational costs.

arXiv AI
3d ago

Harness-Aware Distillation for Small Language Model Agents

The paper introduces Harness-Aware Distillation (HAD), a method for training smaller language model agents that preserves the surrounding harness—software managing context, tools, and feedback—while focusing distillation on the teacher’s contributions beyond the harness. HAD combines an action preference that contrasts teacher actions with and without harness information, and a validity check that filters out contradictory preference pairs. Experiments on long-horizon agent benchmarks show that HAD outperforms standard on‑policy distillation, reducing unproductive loops and improving error recovery without requiring task rewards or future information.

By Moonseok Choi, Taehong Moon, Giung Nam, Juho Lee
Hugging Face Trending Papers
Jul 2

Neuron-Aware Data Selection for Annotation-Free LLM Self-Distillation

Post-training large language models (LLMs) without real-world interaction feedback or human-labeled supervision remains challenging, particularly in specialized domains where expert annotations are costly to obtain. Recent annotation-free self-evolution methods address this by using the model's own outputs as supervision signals, constructing a teacher via additional context and aggregating predictions across multiple rollouts through majority voting to produce pseudo-labels.

Hugging Face Trending Papers
Aug 13

Latent On-Policy Self-Distillation

Enabling agents to learn from experience and internalize it into their policy has become a central problem in self-evolving AI. On-policy self-distillation (OPSD) offers an effective pathway by using a privileged self-teacher to provide dense supervision on the student's own trajectories; however, existing methods still rely heavily on designer-specified privileged artifacts (e.