arXiv AI By Xuanchen Li, Haitao Li, Yujia Zhou, Qingyi Pan, Heng Wang, Yiqun Liu, Min Zhang, Qingyao Ai

Different Feedback, Different Updates: Selective Self-Learning from User Interactions for Large Language Models

Read the original on arXiv AI →

arXiv:2608. 09109v1 Announce Type: new Abstract: User feedback offers natural supervision for persistent LLM improvement, but a single message may support multiple behavioral changes with different scopes of generalization.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jul 2

Neuron-Aware Data Selection for Annotation-Free LLM Self-Distillation

Post-training large language models (LLMs) without real-world interaction feedback or human-labeled supervision remains challenging, particularly in specialized domains where expert annotations are costly to obtain. Recent annotation-free self-evolution methods address this by using the model's own outputs as supervision signals, constructing a teacher via additional context and aggregating predictions across multiple rollouts through majority voting to produce pseudo-labels.

arXiv Machine Learning
Sep 3

Prompt-Space Meta-Learning Does Not Transfer Across Users: A Frozen-LLM Negative Result

The paper investigates whether a frozen large language model can be personalized to individual users via prompt-space meta‑learning. Using the Muse framework, the authors evolve a shared adaptation prompt across a meta‑train user population and test it zero‑shot on over 200 held‑out users in two personalization benchmarks (LaMP‑2 and LaMP‑3). The results show that Muse does not outperform its un‑evolved seed prompt or a control that trains on mismatched user‑support pairs, and it is outperformed by simple few‑shot retrieval on the rating task. The authors attribute this failure to a meta‑objective collapse, where the validation objective is invariant to genuine user‑support correspondence, leading to over‑optimization of instruction polish rather than transferable adaptation.

By Liam Byrne, David Dylan, Orla Fitzgerald, Eoin Doyle, Ciara Nolan, Padraig Lynch, Sinead Gallagher
arXiv AI
Sep 12

Beyond Verified Answers: Solver-Informed Self-Distillation for Bootstrapping Operations Research Language Models

The paper introduces SOLID, a framework that enables operations research language models to self-improve without relying on verified answers or external evaluators. SOLID uses solver-generated artifacts from the model’s own rollouts to create pseudo-references, clustering objectives and applying group-relative advantages for dense self-supervision. Experiments on multiple OR benchmarks show that SOLID enhances solution accuracy for both general-purpose and OR-tuned models compared to outcome-only training.

By Rui Zhu, Minglong Cao, Chenyu Zhou, Jianghao Lin, Dongdong Ge