arXiv AI By Haibo Jin, Xinjie Li, Najmeh Sadoughi, Yang Liu, Yibo Wang, Zhu Liu, Yuzong Liu

Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Aug 28

Beyond Factual QA: Mentorship-Oriented Question Answering over Long-Form Multilingual Content

The paper introduces MentorQA, a multilingual dataset and evaluation framework for mentorship-oriented question answering derived from long‑form videos. It contains nearly 9,000 QA pairs across four languages and defines evaluation dimensions such as clarity, alignment, and learning value that extend beyond factual accuracy. Experiments show that Multi‑Agent QA pipelines outperform other architectures, especially on complex topics and low‑resource languages, while automated LLM‑based evaluation shows variable alignment with human judgments.

By Parth Bhalerao, Diola Dsouza, Ruiwen Guan, Oana Ignat
arXiv AI
Sep 2

WorldBench: Culturally Grounded Benchmark for Multilingual Agents

WorldBench is a new multilingual benchmark that tests large language model agents on culturally grounded everyday workflows, offering 1,600 tasks in seven languages and eight cultures. The benchmark evaluates agents through structured sandbox actions and introduces Constrained Task Success (CTS), a metric that assesses task completion, minimal modification, and other complementary aspects via deterministic and LLM-as-a-Judge evaluations. Experiments show that even leading models achieve only 49.2% CTS, revealing significant gaps in correctness and state preservation across languages and cultures.

By Leonardo Ranaldi, Sherrie Shen, Jushi Kai, Alexandra Birch
arXiv AI
Sep 4

Translation as a Decision Space: A Multi-Agent Perspective on Low-Resource Dialect Generation

The paper proposes treating translation as a structured decision space explored by multiple autonomous agents, rather than producing a single output. Using Turkish–Syrian Arabic dialogue, three agents—zero‑shot, dialect‑stabilized, and pivot translation—are compared on 5,000 sentences, with stabilization nearly doubling dialect marker usage and reducing structural instability. The study introduces an interpretability framework that quantifies decision flexibility through dialect marker frequency, lexical proximity, and structural variance.

By Hasan Alkhder, Mohammad Abboush, Igor Tchappi, Ahmet Zengin, Amro Najjar