arXiv Machine Learning By Daekeun Kim

Naver-News-KO: A Korean News Summarization Dataset for Open-Source Fine-Tuning of Summarization Models

Read the original on arXiv Machine Learning →

arXiv:2607. 20442v1 Announce Type: cross Abstract: We release Naver-News-KO, a Korean news summarization dataset of 27,400 (document, summary) pairs collected from Naver News over a ten-day window in July 2022 across two categories (Economy and IT/Science; 77/23 split), with train/validation/test partitions of 22,194 / 2,466 / 2,740 and a mean per-record document-to-summary character-compression ratio of 6.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 24

PSK at WMT 2026 MIST: Task-Specialized QLoRA Adapters for Multilingual Summarization and Question Answering

The PSK submission to the WMT 2026 Multilingual Instruction Shared Task employs a 3.35B‑parameter Tiny Aya Global model enhanced with three QLoRA adapters, each dedicated to a specific task: multilingual summarization, passage‑based question answering, and filtered standalone question answering. The summarization adapter is trained on multilingual document‑summary pairs, including scientific papers with author‑written abstracts, and outperforms a multitask adapter trained solely on organizer data on a held‑out split. For open question answering, results vary with answer length and evaluation method, prompting the submission of three systems that share the same context and summarization adapters but differ in their open‑QA adapters.

By Srikar Kashyap Pulipaka
arXiv AI
Sep 25

Can Classical Semantic-Extractive Summarization Be Evaluated in Hindi? A Replication Study

The study replicates a distributional‑semantics extractive summarisation method for Hindi, adapting all language‑specific components to Devanagari. Evaluated on the Hindi portions of XL‑Sum and FIRE ILSUM 2.0 with a Devanagari‑aware ROUGE scorer, the replicated system performs significantly worse than a simple three‑sentence lead baseline. Feature ablation shows that sentence position alone reproduces the lead baseline, while other features only steer extraction toward long, entity‑dense body sentences, and TextRank performs identically. "whyItMatters":"The results indicate that current Hindi summarisation benchmarks cannot reward non‑lead content selection, highlighting the need for purpose‑built evaluation resources."

By Showket Ahmad Khan, Mudasir Mohd, Nasrullah Sheikh, Mohsin Altaf Wani, Abid Hussain Wani, Hilal Ahmad Khanday, Niyaz Ahmad Wani