arXiv AI By Sami Shames El Deen, Mariette Awad

Extractive Summarization for Arabic Documents Using SAraBERT with a Semantic Siamese Similarity Evaluation Metric

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Sep 3

Evaluating the Evaluator: Summarization Metrics and LLM-Judges beyond English

The paper introduces BASSE, a multilingual meta‑evaluation dataset containing 2,040 human‑rated abstractive summaries produced manually or by five LLMs with four prompts. Annotators scored each summary on coherence, consistency, fluency, relevance, and 5W1H using a 5‑point Likert scale. Benchmarking shows proprietary LLM‑judge models best align with human judgments, followed by criteria‑specific automatic metrics, while open‑source judge LLMs perform poorly.

By Jeremy Barnes, Naiara Perez, Alba Bonet-Jover, Bego\~na Altuna
arXiv Machine Learning
Sep 11

E-CONAN (Entailment, CONtradition And Neutral) Benchmarks: Arabic Textual Entailment and Natural Inference Datasets

E-CONAN introduces Arabic textual entailment and natural inference benchmarks comprising two datasets: E-CONAN-2 (2-way RTE) and E-CONAN-3 (3-way NLI). The datasets are built from automatically-translated pairs, human-validated machine translations, hand-crafted pairs from Arabic teaching books, and rumor-containing news headlines. The authors evaluated nine multilingual pretrained models and five large language models on these benchmarks, demonstrating that E-CONAN offers a more diverse and robust assessment than existing datasets like XNLI and ArNLI.

By Khloud AL Jallad, Nada Ghneim, Ghaida Rebdawi