arXiv:2607. 02460v1 Announce Type: cross Abstract: Post-training large language models (LLMs) without real-world interaction feedback or human-labeled supervision remains challenging, particularly in specialized domains where expert annotations are costly to obtain.
By Zhuowei Chen, Xiang Lorraine Li
Post-training large language models (LLMs) without real-world interaction feedback or human-labeled supervision remains challenging, particularly in specialized domains where expert annotations are costly to obtain. Recent annotation-free self-evolution methods address this by using the model's own outputs as supervision signals, constructing a teacher via additional context and aggregating predictions across multiple rollouts through majority voting to produce pseudo-labels.
arXiv:2606. 27634v1 Announce Type: new Abstract: Small Language Models (SLMs) are increasingly being considered for deployment on edge devices such as laptops, enabling private, low-latency, and locally personalized applications.
By Thomas S. Paula, Lucas S. Kupssinsk\"u, Rodrigo C. Barros
The paper investigates whether a frozen large language model can be personalized to individual users via prompt-space meta‑learning. Using the Muse framework, the authors evolve a shared adaptation prompt across a meta‑train user population and test it zero‑shot on over 200 held‑out users in two personalization benchmarks (LaMP‑2 and LaMP‑3). The results show that Muse does not outperform its un‑evolved seed prompt or a control that trains on mismatched user‑support pairs, and it is outperformed by simple few‑shot retrieval on the rating task. The authors attribute this failure to a meta‑objective collapse, where the validation objective is invariant to genuine user‑support correspondence, leading to over‑optimization of instruction polish rather than transferable adaptation.
By Liam Byrne, David Dylan, Orla Fitzgerald, Eoin Doyle, Ciara Nolan, Padraig Lynch, Sinead Gallagher
The paper introduces SOLID, a framework that enables operations research language models to self-improve without relying on verified answers or external evaluators. SOLID uses solver-generated artifacts from the model’s own rollouts to create pseudo-references, clustering objectives and applying group-relative advantages for dense self-supervision. Experiments on multiple OR benchmarks show that SOLID enhances solution accuracy for both general-purpose and OR-tuned models compared to outcome-only training.
By Rui Zhu, Minglong Cao, Chenyu Zhou, Jianghao Lin, Dongdong Ge
arXiv:2608. 05161v1 Announce Type: cross Abstract: Instruction-tuned LLMs are deployed into environments where domains evolve, yet extending a fine-tuned model's capabilities without full retraining remains an unsolved practical challenge.
By Josh McGiff, Salma Mekaoui, Robert Shanahan, Nikola S. Nikolov