arXiv:2608.29884v1 Announce Type: new
Abstract: We show that sequence-level distillation from a capable long-context teacher model is a simple, annotation-free, and data-efficient strategy for improv...
By Mohamed Elaraby, Ahmed Elhady, Diane Litman
arXiv:2406. 14657v4 Announce Type: replace-cross Abstract: We introduce OpenDebateEvidence, a comprehensive dataset for argument mining and summarization sourced from the American Competitive Debate community.
By Allen Roush, Yusuf Shabazz, Arvind Balaji, Peter Zhang, Stefano Mezza, Markus Zhang, Sanjay Basu, Sriram Vishwanath, Mehdi Fatemi, Ravid Shwartz-Ziv
The paper introduces Gavel, a framework for evaluating large language models (LLMs) on long-context legal summarization tasks. Gavel includes a reference-based component (Gavel-Ref) with checklist, residual-fact, and writing-style checks, and a reference-free component (Gavel-Agent) that assesses factual coverage directly from source documents. Experiments on 12 frontier LLMs reveal that models tend to omit key information more than hallucinate, perform well on simple checklist items but struggle with rare, complex items, and their performance degrades with longer cases. Gavel-Agent cuts token usage by at least 36% compared to traditional methods while maintaining competitive accuracy, and it also generalizes effectively to the medical domain.
By Yao Dou, Benjamin Mamut, Wei Xu
arXiv:2608. 04307v1 Announce Type: cross Abstract: Text summarization is deceptively difficult.
By Karen Lee, Dhanashree Balaram, Seojun Shon, Umair Rasheed
arXiv:2606. 03867v1 Announce Type: cross Abstract: Multi-Document Summarization (MDS) plays a critical role in distilling essential information from collections of textual data.
By Cuong Vuong Tuan, Trang Mai Xuan, Tien-Cuong Nguyen, Vu-Duc Ngo, Thien Van Luong
The paper introduces CARPAS, a new task that dynamically refines user-provided aspects for aspect-based summarization in large language models (LLMs). It presents three new datasets and evaluates four prompting strategies, finding that LLMs tend to over-generate aspects, leading to overly long and misaligned summaries. To address this, the authors propose a two-stage framework that first generates lightweight scope guidance before aspect refinement and summarization, which improves focus, reduces over-generation, and enhances performance across all datasets.
By Yong-En Tian, Yu-Chien Tang, An-Zi Yen, Wen-Chih Peng
The paper introduces BASSE, a multilingual meta‑evaluation dataset containing 2,040 human‑rated abstractive summaries produced manually or by five LLMs with four prompts. Annotators scored each summary on coherence, consistency, fluency, relevance, and 5W1H using a 5‑point Likert scale. Benchmarking shows proprietary LLM‑judge models best align with human judgments, followed by criteria‑specific automatic metrics, while open‑source judge LLMs perform poorly.
By Jeremy Barnes, Naiara Perez, Alba Bonet-Jover, Bego\~na Altuna
arXiv:2606. 01736v1 Announce Type: cross Abstract: As LLMs are increasingly used to draft public-facing arguments, they may flatten public debate by repeatedly introducing the same polished, plausible arguments.
By Yekyung Kim, Yapei Chang, Chau Minh Pham, Mohit Iyyer
arXiv:2606. 08000v1 Announce Type: cross Abstract: The progress of large language models (LLMs) has fueled claims that model-generated summaries rival or even surpass human-written references, raising questions about whether summarization remains an open research problem.
By Dongqi Liu, Chenxi Whitehouse, Zheng Zhao, Zhuchen Cao, Jian Li, Yabiao Wang
Scientific long-document summarization datasets commonly treat author-written abstracts as gold reference summaries, although their quality and alignment with the source article vary. At the same time, publicly available scientific summarization datasets remain limited in scale and structure for modern long-context models.
arXiv:2608. 03655v1 Announce Type: cross Abstract: Abstractive summarization models remain vulnerable to factual inconsistency, redundancy, and weak length control.
By Zeyu Wang, Guanghua Wang, Meng Xu
CLASE is a hybrid evaluation method for Chinese legal text that combines linguistic feature-based scores with experience-guided LLM-as-a-judge scores. It learns from contrastive pairs of authentic legal documents and their LLM-generated counterparts, enabling transparent, reference-free assessment of stylistic quality. Experiments on 200 Chinese legal documents show that CLASE aligns better with human judgments than traditional metrics and offers interpretable score breakdowns and improvement suggestions.
By Yiran Rex Ma, Yuxiao Ye, Huiyuan Xie