Learning to summarize with human feedback
We’ve applied reinforcement learning from human feedback to train language models that are better at summarization.
We trained “critique-writing” models to describe flaws in summaries. Human evaluators find flaws in summaries much more often when shown our model’s critiques.
We’ve applied reinforcement learning from human feedback to train language models that are better at summarization.
Scaling human oversight of AI systems for tasks that are difficult to evaluate.
arXiv:2605. 03202v2 Announce Type: replace Abstract: Large language models offer a tempting solution to address the peer review crisis.
arXiv:2606. 08000v1 Announce Type: cross Abstract: The progress of large language models (LLMs) has fueled claims that model-generated summaries rival or even surpass human-written references, raising questions about whether summarization remains an open research problem.
arXiv:2601. 14171v2 Announce Type: replace Abstract: Writing effective rebuttals is a high-stakes task that demands more than linguistic fluency, as it requires precise alignment between reviewer intent and manuscript details.
arXiv:2510. 26518v2 Announce Type: replace Abstract: Human feedback is critical for aligning AI systems to human values.
arXiv:2606. 05436v1 Announce Type: new Abstract: Summarizing the latest medical literature to guide clinical decision-making is essential for evidence-based medicine and high-quality patient care.
arXiv:2605. 28882v2 Announce Type: replace-cross Abstract: With the rapid advancement of large language models, evaluating human-likeness in open-ended conversation has become increasingly important.
arXiv:2606. 06481v1 Announce Type: cross Abstract: As AI writing assistants become increasingly integrated into real-world drafting and revision workflows, many documents are no longer purely human-written or AI-generated, but instead result from progressive human-AI co-editing.
arXiv:2607. 10390v1 Announce Type: cross Abstract: Although large language models (LLMs) have shown promising potential in news summarization tasks, their performance on long-document summarization remains challenging as their length often exceeds the input limits.
arXiv:2605. 17064v2 Announce Type: replace Abstract: Large language models are optimized for instruction following and agentic tasks remain poorly aligned with the requirements of high-quality creative writing.
arXiv:2604. 21827v2 Announce Type: replace Abstract: In accomplishing complex tasks, human cognition typically progresses from abstract to concrete (e.