arXiv:2507. 06506v2 Announce Type: replace-cross Abstract: Translating wordplay across languages presents unique challenges that have long confounded both professional human translators and machine translation systems.
By Russell Taylor, Benjamin Herbert, Michael Sana
Internet memes are a pervasive form of multimodal online communication; however, such communication often involves users from diverse linguistic and cultural backgrounds. Therefore, adapting memes acr...
TransMeme introduces a multi‑agent framework for cross‑cultural meme transcreation, addressing the unique challenges of preserving intent, adapting cultural meaning, and maintaining multimodal consistency. The system coordinates specialized agents for cultural adaptation, text rewriting, revision, and visual adjustment, and is evaluated on Chinese‑English meme pairs. Human and LLM‑based evaluations show that TransMeme outperforms baselines, achieving a 33.1% average improvement in human scores and a 60% Top‑1 ranking rate in LLM judgments.
By Jingyi Zheng, Yule Liu, Zifan Peng, Tianyi Hu, Yuemeng Zhao, Xinhu Zheng, Xinlei He
arXiv:2609.23490v1 Announce Type: new
Abstract: Large language model (LLM) agents increasingly execute multi-step workflows through tool use and interaction with users and environments. However, curr...
By Peng Kuang, Yuchun Fan, Jiangnan Li, Minghao Wu, Jialong Tang, Hao-Ran Wei, Weixuan Wang, Jianhong Tu, Baosong Yang, Tong Xiao
arXiv:2609.38660v1 Announce Type: cross
Abstract: Long-form subtitle translation requires reasoning over discourse and cultural context spanning episodes or entire series, while maintaining consisten...
By Haibo Jin, Xinjie Li, Najmeh Sadoughi, Yang Liu, Yibo Wang, Zhu Liu, Yuzong Liu
Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost exclusively in English. As AI agents are globally deployed to a linguistically diverse user base, whether agentic competence measured in English transfers to other languages remains an open question.
arXiv:2607. 13189v1 Announce Type: cross Abstract: We present RAGthoven, our system for SemEval-2026 Task 1 (MWAHAHA), Subtask A (multilingual constrained humor generation in English, Spanish, and Chinese).
By Marek \v{S}uppa, Vikt\'oria Ondrejov\'a, Lucia Ganajov\'a, Gregor Karetka, Daniel Skala
The paper proposes treating translation as a structured decision space explored by multiple autonomous agents, rather than producing a single output. Using Turkish–Syrian Arabic dialogue, three agents—zero‑shot, dialect‑stabilized, and pivot translation—are compared on 5,000 sentences, with stabilization nearly doubling dialect marker usage and reducing structural instability. The study introduces an interpretability framework that quantifies decision flexibility through dialect marker frequency, lexical proximity, and structural variance.
By Hasan Alkhder, Mohammad Abboush, Igor Tchappi, Ahmet Zengin, Amro Najjar
arXiv:2608. 10444v1 Announce Type: cross Abstract: Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chains.
By Si'an Xie (Beijing University of Posts and Telecommunications), Jiaxun Liu (Peking University), Biao Yang (Kuaishou Technology), Wei Yuan (Kuaishou Technology), Fan Yang (Kuaishou Technology), Tingting Gao (Kuaishou Technology), Ming Wu (Beijing University of Posts and Telecommunications)
arXiv:2606.28715v2 Announce Type: replace-cross
Abstract: While AI development and evaluation for Southeast Asia (SEA) has grown rapidly, agent capabilities in regional languages are still poorly und...
By My Chiffon Nguyen, Aulia Adila, Saksorn Ruangtanusak, Kittiphat Leesombatwathana, Vissuta Gunawan Lim, Patomporn Payoungkhamdee, Samuel Cahyawijaya
EviSI is an evidence‑based evaluation agent for low‑latency simultaneous speech‑to‑speech translation. It combines Multidimensional Quality Metrics with interpreter‑developed criteria, using shared source evidence to assess four dimensions—Anchor, Event, Logic, and Fluency—while deduplicating verified errors before scoring. On English‑to‑Chinese and Chinese‑to‑English data, EviSI’s rankings correlate strongly with human judgments, outperforming BLEU and COMET, and its multilingual extension maintains these correlations across five language directions.
By Ben Yan, Zongyao Li, Xiaoyu Chen, Daimeng Wei, Weidong Liu, Huan Zhao, Chong Li, Yaode Wang, Yuzhe Shang
The paper introduces KVoiceBench, KOpenAudioBench, and KMMAU—three Korean speech benchmarks created through agent-driven frameworks that adapt existing SpokenQA and ASR resources into Korean SpokenQA and audio understanding tasks. These benchmarks total 12,345 samples and are publicly released to evaluate SpeechLMs beyond English. The authors benchmark eight recent SpeechLMs, revealing significant English‑Korean performance gaps and divergent rankings between SpokenQA and audio understanding, highlighting multilingual weaknesses not apparent in English-only tests.
By Haechan Kim, Seungjun Chung, Inkyu Park, Jihoo Lee, Jonghyun Lee