arXiv AI

Different Feedback, Different Updates: Selective Self-Learning from User Interactions for Large Language Models

arXiv:2608. 09109v1 Announce Type: new Abstract: User feedback offers natural supervision for persistent LLM improvement, but a single message may support multiple behavioral changes with different scopes of generalization.

Hugging Face Trending Papers
Jul 2

Neuron-Aware Data Selection for Annotation-Free LLM Self-Distillation

Post-training large language models (LLMs) without real-world interaction feedback or human-labeled supervision remains challenging, particularly in specialized domains where expert annotations are costly to obtain. Recent annotation-free self-evolution methods address this by using the model's own outputs as supervision signals, constructing a teacher via additional context and aggregating predictions across multiple rollouts through majority voting to produce pseudo-labels.

arXiv Machine Learning
Sep 3

Prompt-Space Meta-Learning Does Not Transfer Across Users: A Frozen-LLM Negative Result

The paper investigates whether a frozen large language model can be personalized to individual users via prompt-space meta‑learning. Using the Muse framework, the authors evolve a shared adaptation prompt across a meta‑train user population and test it zero‑shot on over 200 held‑out users in two personalization benchmarks (LaMP‑2 and LaMP‑3). The results show that Muse does not outperform its un‑evolved seed prompt or a control that trains on mismatched user‑support pairs, and it is outperformed by simple few‑shot retrieval on the rating task. The authors attribute this failure to a meta‑objective collapse, where the validation objective is invariant to genuine user‑support correspondence, leading to over‑optimization of instruction polish rather than transferable adaptation.

By Liam Byrne, David Dylan, Orla Fitzgerald, Eoin Doyle, Ciara Nolan, Padraig Lynch, Sinead Gallagher
arXiv AI
Sep 12

Beyond Verified Answers: Solver-Informed Self-Distillation for Bootstrapping Operations Research Language Models

The paper introduces SOLID, a framework that enables operations research language models to self-improve without relying on verified answers or external evaluators. SOLID uses solver-generated artifacts from the model’s own rollouts to create pseudo-references, clustering objectives and applying group-relative advantages for dense self-supervision. Experiments on multiple OR benchmarks show that SOLID enhances solution accuracy for both general-purpose and OR-tuned models compared to outcome-only training.

By Rui Zhu, Minglong Cao, Chenyu Zhou, Jianghao Lin, Dongdong Ge
arXiv AI
Jul 1

What Drives Interactive Improvement from Feedback?

arXiv:2606. 30774v1 Announce Type: new Abstract: We study when natural-language feedback produces improvement beyond the gains obtainable from repeated attempts alone.

By Bart{\l}omiej Cupia{\l}, Jan {\L}ojek, Miko{\l}aj Garstecki, Szymon Pob{\l}ocki, Alicja Ziarko, Piotr Mi{\l}o\'s
arXiv AI
Sep 24

Agent-Editing World Model: Rethinking World Modeling for LLM Agents

The paper introduces the Agent-Editing World Model (AEWM), a new approach that models how reasoning and actions influence future task progress instead of simulating tool responses. AEWM includes an Action Judge that classifies decisions as Critical, Exploratory, or Noisy, and a State Revision mechanism that edits noisy reasoning–action continuations from the same observed history. The integrated system, EditAct, directly updates the underlying state during real execution, leading to significant performance gains across multiple benchmarks and agent backbones.

By Shuang Sun, Guoxin Chen, Fanzhe Meng, Jia Deng, Huatong Song, Jinhao Jiang, Wayne Xin Zhao, Hongteng Xu, Ji-Rong Wen
arXiv AI
3d ago

ConflictGuide: AutoResearch Improves When Competing Behaviors Are Made Visible

The paper introduces ConflictGuide, a method that enhances LLM-based AutoResearch by incorporating feedback on competing behaviors during model code editing. By first exploring with scalar task performance and then using probes to measure and alleviate conflicts, ConflictGuide increases the proportion of edits that improve multiple behaviors and sustains progress beyond scalar-only plateaus. Experiments across five model families show reductions in task and conflict-related errors by up to 28% and 14% compared to scalar-only AutoResearch.

By Binqian Xu, Qiran Zou, Xiangbo Shu, Dianbo Liu