Learning to summarize with human feedback
We’ve applied reinforcement learning from human feedback to train language models that are better at summarization.
We trained “critique-writing” models to describe flaws in summaries. Human evaluators find flaws in summaries much more often when shown our model’s critiques.
We’ve applied reinforcement learning from human feedback to train language models that are better at summarization.
The study evaluates AI-generated summaries for cancer patients using a dual assessment framework that includes human experts and LLM-as-a-judge. Human domain experts—oncology clinicians and patient-facing care staff—assess summary quality on accuracy, clinical relevance, and readability. The research identifies limitations such as omissions and minor inaccuracies, which are then used to iteratively refine prompts, grounding, and safety guardrails.
arXiv:2609.14738v1 Announce Type: new Abstract: Automated reviewing systems are increasingly evaluated based on the quality of the reviews they produce. Yet a review is only useful if acting on it le...
Scaling human oversight of AI systems for tasks that are difficult to evaluate.
arXiv:2605. 03202v2 Announce Type: replace Abstract: Large language models offer a tempting solution to address the peer review crisis.
arXiv:2606. 08000v1 Announce Type: cross Abstract: The progress of large language models (LLMs) has fueled claims that model-generated summaries rival or even surpass human-written references, raising questions about whether summarization remains an open research problem.
arXiv:2608.30842v1 Announce Type: new Abstract: Humans play a vital role at every stage of AI development, from data collection and curation to model development and evaluation. However, humans often...
arXiv:2601. 14171v2 Announce Type: replace Abstract: Writing effective rebuttals is a high-stakes task that demands more than linguistic fluency, as it requires precise alignment between reviewer intent and manuscript details.
arXiv:2510. 26518v2 Announce Type: replace Abstract: Human feedback is critical for aligning AI systems to human values.
The paper examines whether existing automatic methods can reliably assess creativity in text produced by large language models (LLMs). By collecting human ratings on 11 creativity dimensions for both human and AI short stories, the authors compare these judgments with automated metrics and LLM-as-a-Judge evaluations. The results show a significant misalignment: automated metrics and LLM judges favor AI-generated stories and show near-zero correlation with human assessments, revealing fundamental limitations in current computational approaches to evaluating creative text.
arXiv:2606. 05436v1 Announce Type: new Abstract: Summarizing the latest medical literature to guide clinical decision-making is essential for evidence-based medicine and high-quality patient care.
arXiv:2605. 28882v2 Announce Type: replace-cross Abstract: With the rapid advancement of large language models, evaluating human-likeness in open-ended conversation has become increasingly important.