arXiv:2606. 05494v3 Announce Type: replace-cross Abstract: Automatic text summarization has become increasingly important due to the rapid growth of digital textual information.
By Ahmed Alansary, Ali Hamdi
arXiv:2606. 05494v1 Announce Type: cross Abstract: Automatic text summarization has become increasingly important due to the rapid growth of digital textual information.
By Ahmed Alansary, Ali Hamdi
arXiv:2607. 10806v1 Announce Type: cross Abstract: Quantifying abstractiveness in generated summaries is essential for evaluating summarization models beyond surface-level metrics like ROUGE.
By Praveenkumar Katwe, Rakesh Chandra Balabantaray, Kali Prasad Vittala
The paper introduces BASSE, a multilingual meta‑evaluation dataset containing 2,040 human‑rated abstractive summaries produced manually or by five LLMs with four prompts. Annotators scored each summary on coherence, consistency, fluency, relevance, and 5W1H using a 5‑point Likert scale. Benchmarking shows proprietary LLM‑judge models best align with human judgments, followed by criteria‑specific automatic metrics, while open‑source judge LLMs perform poorly.
By Jeremy Barnes, Naiara Perez, Alba Bonet-Jover, Bego\~na Altuna
arXiv:2608. 19200v1 Announce Type: cross Abstract: Text summarization refers to the task of condensing a document into a shorter version while preserving its key information.
By Daisy Aptovska, Vinayak Elangovan
E-CONAN introduces Arabic textual entailment and natural inference benchmarks comprising two datasets: E-CONAN-2 (2-way RTE) and E-CONAN-3 (3-way NLI). The datasets are built from automatically-translated pairs, human-validated machine translations, hand-crafted pairs from Arabic teaching books, and rumor-containing news headlines. The authors evaluated nine multilingual pretrained models and five large language models on these benchmarks, demonstrating that E-CONAN offers a more diverse and robust assessment than existing datasets like XNLI and ArNLI.
By Khloud AL Jallad, Nada Ghneim, Ghaida Rebdawi
arXiv:2606. 19591v1 Announce Type: cross Abstract: In this technical report, we focus on solving the challenge of Vietnamese multi-document abstractive summarization, introduced in the International Workshop on Vietnamese Language and Speech Processing (VLSP) 2022.
By Vu Nguyen Nguyen Xuan, Huy Ngo Quang
The paper introduces a Nepali Question‑Answer dataset focused on passport‑related FAQs to support information retrieval in a low‑resource language. The authors fine‑tune transformer‑based embedding models for semantic similarity and compare them against the BM25 baseline. Their experiments show that fine‑tuned SBERT models outperform BM25, while multilingual E5 embeddings achieve the best overall retrieval performance.
By Funghang Limbu Begha, Praveen Acharya, Bal Krishna Bal
Arabic Natural Language Processing (NLP) has grown rapidly over the past decade, driven by digital transformation in the Arab world, social media, and large language models (LLMs). Despite this growth...
arXiv:2607. 21010v1 Announce Type: new Abstract: Zero-shot summarization using Large Language Models (LLMs) has significantly advanced the abstractive summarization task by producing coherent and fluent summaries.
By Vasudha Bhatnagar, Purnima Bindal, Vikas Kumar, Raj Kumari Bahl
The paper proposes a new approach to single-document extractive summarization by constructing a sentence hypergraph where sentences are nodes and keywords or named entities are hyperedges. A greedy algorithm is then used to find a dominating set of this hypergraph, which yields the sentences that compose the summary. The study compares this hypergraph-based method with existing graph-based summarization techniques.
By Aamir Miyajiwala, Aabha Pingle, Sheetal Sonawane, Surajit Kr. Nath
arXiv:2608.23421v1 Announce Type: new
Abstract: Natural Language Processing (NLP) has grown rapidly over the past decade, driven by digital transformation in the Arab world, social media, and large l...
By Mullosharaf K. Arabov