Peer-Predictive Self-Training for Language Model Reasoning
arXiv:2604. 13356v3 Announce Type: replace-cross Abstract: Mechanisms for continued self-improvement of language models without external supervision remain an open challenge.
arXiv:2603. 19294v4 Announce Type: replace Abstract: While post-training has successfully improved large language models (LLMs) across a variety of domains, these gains heavily rely on human-labeled data or external verifiers.
arXiv:2604. 13356v3 Announce Type: replace-cross Abstract: Mechanisms for continued self-improvement of language models without external supervision remain an open challenge.
arXiv:2510. 05342v2 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) has emerged as a simple and effective method for aligning large language models.
arXiv:2507.21931v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) often produce plausible but poorly-calibrated answers, limiting their reliability on reasoning-intensive tasks....
arXiv:2608. 09507v1 Announce Type: cross Abstract: Natural language user preferences provide an interpretable interface for LLM personalization.
Task-Level Natural Language Priors as Learning Signals for Low-Resource LLM Training proposes Prior-Guided Tuning (PGT), a training approach that treats natural-language priors as auxiliary learning signals rather than just input context. The method introduces Contrastive Prior Steering (CPS), which adds positive and negative prior-conditioned auxiliary losses while preserving the original supervised objective. Experiments on AmbiMath, Jigsaw, and MNLI/HANS demonstrate that CPS consistently outperforms plain and prompt fine-tuning, achieving high accuracy and significant gains with limited training data.
arXiv:2607. 02460v1 Announce Type: cross Abstract: Post-training large language models (LLMs) without real-world interaction feedback or human-labeled supervision remains challenging, particularly in specialized domains where expert annotations are costly to obtain.
arXiv:2509. 03647v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer from "self-preference bias": a tendency to favor their own outputs over those of other models.
arXiv:2606. 00544v1 Announce Type: new Abstract: Modern language-model fine-tuning typically pairs each prompt with a single response, even though many prompts admit multiple valid completions.
Post-training large language models (LLMs) without real-world interaction feedback or human-labeled supervision remains challenging, particularly in specialized domains where expert annotations are costly to obtain. Recent annotation-free self-evolution methods address this by using the model's own outputs as supervision signals, constructing a teacher via additional context and aggregating predictions across multiple rollouts through majority voting to produce pseudo-labels.
GrammarRL introduces a label‑free reinforcement learning approach that adapts language models to grammar constraints without annotated data. It optimizes two self‑supervised rewards—direct and reverse—using a Reinforce Leave‑One‑Out objective over grammar‑constrained rollouts, and regularizes toward a frozen base model. Experiments on sign‑language gloss translation, hierarchical text classification, and named entity recognition with Llama models show consistent gains over constrained greedy decoding and competitive performance to beam search while keeping inference cost low.
arXiv:2606. 05734v1 Announce Type: new Abstract: Large language models (LLMs) are generally constrained from expressing feelings through human-preference alignment in post-training processes.
arXiv:2609.15972v1 Announce Type: cross Abstract: As language models become more capable, long-term collaboration in learning, reasoning, and decision-making calls for a deeper understanding of the p...