arXiv:2606. 28715v1 Announce Type: cross Abstract: While AI development and evaluation for Southeast Asia (SEA) has grown rapidly, agent capabilities in regional languages are still poorly understood despite its importance to sovereign AI.
By My Chiffon Nguyen, Aulia Adila, Saksorn Ruangtanusak, Kittiphat Leesombatwathana, Vissuta Gunawan Lim, Patomporn Payoungkhamdee, Samuel Cahyawijaya
arXiv:2609.24812v1 Announce Type: new
Abstract: Voice provides a natural and immediate interface for AI agents. Many settings in which voice agents could be useful, including meetings, households, an...
By Chenxu Xiong, Dongming Shen, Yuzhi Tang, Wentao Ma, Mu Li, Alex Smola
arXiv:2606.28715v2 Announce Type: replace-cross
Abstract: While AI development and evaluation for Southeast Asia (SEA) has grown rapidly, agent capabilities in regional languages are still poorly und...
By My Chiffon Nguyen, Aulia Adila, Saksorn Ruangtanusak, Kittiphat Leesombatwathana, Vissuta Gunawan Lim, Patomporn Payoungkhamdee, Samuel Cahyawijaya
The paper introduces KVoiceBench, KOpenAudioBench, and KMMAU—three Korean speech benchmarks created through agent-driven frameworks that adapt existing SpokenQA and ASR resources into Korean SpokenQA and audio understanding tasks. These benchmarks total 12,345 samples and are publicly released to evaluate SpeechLMs beyond English. The authors benchmark eight recent SpeechLMs, revealing significant English‑Korean performance gaps and divergent rankings between SpokenQA and audio understanding, highlighting multilingual weaknesses not apparent in English-only tests.
By Haechan Kim, Seungjun Chung, Inkyu Park, Jihoo Lee, Jonghyun Lee
arXiv:2609.38867v1 Announce Type: new
Abstract: Large language model (LLM) computer-use agents are typically evaluated with clean written instructions, despite speech being an increasingly popular in...
By Terumi Chiba, Guangzhi Sun, Zheqi Yuan, Chao Zhang
arXiv:2608.28641v1 Announce Type: cross
Abstract: Most evaluations for coding agents are conducted exclusively in English, which does not reflect real-world multilingual deployment. We present Termin...
By Yunsu Kim, Kaden Uhlig, Ashwin Purohit, Milind Agarwal, Patrick Simianer, Anil Arslan, Kiarash Mokhtari, Thomas Zenkel, Johannes Mosig, Gabriel Bretschner, Shamik Bose, Joern Wuebker, John DeNero
arXiv:2609.23490v1 Announce Type: new
Abstract: Large language model (LLM) agents increasingly execute multi-step workflows through tool use and interaction with users and environments. However, curr...
By Peng Kuang, Yuchun Fan, Jiangnan Li, Minghao Wu, Jialong Tang, Hao-Ran Wei, Weixuan Wang, Jianhong Tu, Baosong Yang, Tong Xiao
arXiv:2601. 05366v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly deployed as agents that invoke external tools through structured function calls.
By Zheng Luo, T Pranav Kutralingam, Ogochukwu N Okoani, Wanpeng Xu, Hua Wei, Xiyang Hu
SEA-SpeechBench is a large‑scale multitask benchmark for speech understanding in 11 Southeast Asian languages, comprising 97,194 samples across 99 evaluation sets and 597 hours of curated audio. It covers nine tasks in three categories—speech processing, paralinguistic analysis, and a novel temporal understanding dimension—using multilingual prompting in both native SEA languages and English. Evaluation of current models shows significant performance gaps, especially in temporal understanding, emotion recognition, and speech translation, with low‑resource languages lagging behind English by up to 41 percentage points.
By Jingyi Liao, Wenyu Zhang, Zhuohan Liu, Yingxu He, Geyu Lin, Xunlong Zou, Shuo Sun, Syed Ali Redha Alsagoff, Ai Ti Aw
WorldBench is a new multilingual benchmark that tests large language model agents on culturally grounded everyday workflows, offering 1,600 tasks in seven languages and eight cultures. The benchmark evaluates agents through structured sandbox actions and introduces Constrained Task Success (CTS), a metric that assesses task completion, minimal modification, and other complementary aspects via deterministic and LLM-as-a-Judge evaluations. Experiments show that even leading models achieve only 49.2% CTS, revealing significant gaps in correctness and state preservation across languages and cultures.
By Leonardo Ranaldi, Sherrie Shen, Jushi Kai, Alexandra Birch
Despite the growing use of LLM-powered agents to solve multi-step tasks in complex environments, existing benchmarks rarely test state preservation, performance across languages, and application to re...
The rapid advancement of audio and multimodal large language models has unlocked transformative speech understanding capabilities, yet evaluation frameworks remain predominantly English-centric, leavi...