arXiv:2608. 19200v1 Announce Type: cross Abstract: Text summarization refers to the task of condensing a document into a shorter version while preserving its key information.
By Daisy Aptovska, Vinayak Elangovan
The PSK submission to the WMT 2026 Multilingual Instruction Shared Task employs a 3.35B‑parameter Tiny Aya Global model enhanced with three QLoRA adapters, each dedicated to a specific task: multilingual summarization, passage‑based question answering, and filtered standalone question answering. The summarization adapter is trained on multilingual document‑summary pairs, including scientific papers with author‑written abstracts, and outperforms a multitask adapter trained solely on organizer data on a held‑out split. For open question answering, results vary with answer length and evaluation method, prompting the submission of three systems that share the same context and summarization adapters but differ in their open‑QA adapters.
By Srikar Kashyap Pulipaka
We’ve applied reinforcement learning from human feedback to train language models that are better at summarization.
arXiv:2606. 19591v1 Announce Type: cross Abstract: In this technical report, we focus on solving the challenge of Vietnamese multi-document abstractive summarization, introduced in the International Workshop on Vietnamese Language and Speech Processing (VLSP) 2022.
By Vu Nguyen Nguyen Xuan, Huy Ngo Quang
arXiv:2606. 03867v1 Announce Type: cross Abstract: Multi-Document Summarization (MDS) plays a critical role in distilling essential information from collections of textual data.
By Cuong Vuong Tuan, Trang Mai Xuan, Tien-Cuong Nguyen, Vu-Duc Ngo, Thien Van Luong
arXiv:2608. 04307v1 Announce Type: cross Abstract: Text summarization is deceptively difficult.
By Karen Lee, Dhanashree Balaram, Seojun Shon, Umair Rasheed
arXiv:2606. 08445v1 Announce Type: cross Abstract: Meeting documents are challenging to summarize due to their length and complex conversational structure.
By Sangwon Ryu, Heejin Do, Jun Seo, Daehui Kim, Yunsu Kim, Gary Geunbae Lee, Jungseul Ok
arXiv:2607. 20442v1 Announce Type: cross Abstract: We release Naver-News-KO, a Korean news summarization dataset of 27,400 (document, summary) pairs collected from Naver News over a ten-day window in July 2022 across two categories (Economy and IT/Science; 77/23 split), with train/validation/test partitions of 22,194 / 2,466 / 2,740 and a mean per-record document-to-summary character-compression ratio of 6.
By Daekeun Kim
Loss-Based Active Learning for Neural Abstractive Summarization proposes LOBSTER, an active learning framework that selects unlabeled documents similar to the model’s high‑loss training examples to correct specific weaknesses. The method is tailored for abstractive summarization, addressing instability and computational bottlenecks seen in prior work. Experiments on three benchmark datasets and two backbone models show that LOBSTER matches or surpasses state‑of‑the‑art performance while speeding up query selection by up to 665×.
By Michail Ioannou, Tatiana Passali, George Michalopoulos, Grigorios Tsoumakas
arXiv:2606. 05494v1 Announce Type: cross Abstract: Automatic text summarization has become increasingly important due to the rapid growth of digital textual information.
By Ahmed Alansary, Ali Hamdi
arXiv:2608.29884v1 Announce Type: new
Abstract: We show that sequence-level distillation from a capable long-context teacher model is a simple, annotation-free, and data-efficient strategy for improv...
By Mohamed Elaraby, Ahmed Elhady, Diane Litman
arXiv:2607. 17524v1 Announce Type: cross Abstract: We propose Token-Level Off-Policy Labeling (TOPL), an off-policy training paradigm that reframes post-training as a token-level correctness prediction task.
By Zitong Huang, Gustavo Lucas Carvalho, Deqing Fu, Robin Jia