Large language models are increasingly deployed as general-purpose educational and technical assistance systems, but their underlying infrastructure does not treat languages equally. One underexamined source of disparity is tokenization: semantically equivalent content can require substantially different token counts across languages, affecting API cost, latency, and usable context length before a model is invoked.
arXiv:2607. 06544v1 Announce Type: new Abstract: As Artificial Intelligence (AI) makes inroads into different parts of the Indian subcontinent, there is significant interest in studying how AI impacts the linguistic and cultural foundations of this civilization.
By Aparna Madva, Sharath Srivatsa, Srinath Srinivasa, Tulika Saha
arXiv:2606. 30656v1 Announce Type: cross Abstract: Artificial Intelligence (AI) has the potential to be transformative for development, but Africa is currently facing a fragmented and challenging "AI divide".
By Abayomi O. Agbeyangi, Jose M. Lukose
arXiv:2608. 13695v1 Announce Type: cross Abstract: Large language model providers routinely cite multilingual safety benchmarks spanning a dozen or more languages as evidence that their models are safe for non-English-speaking users.
By Chialuka Prisca-Mary Onuoha, Bright Etornam Sunu, Rashidat Sikiru
arXiv:2606. 26112v1 Announce Type: cross Abstract: Low-resource languages face a critical challenge in AI development: creating specialized conversational systems without access to massive training corpora.
By Siddhant Hitesh Mantri, Dhara Gorasiya, Malhar Kulkarni, Pushpak Bhattacharya
arXiv:2608.30092v1 Announce Type: cross
Abstract: We present Arkios, a 1.04B-parameter dense transformer pretrained from scratch on 150B tokens of bilingual English-Nepali text, using a custom single...
By Sajal Regmi, Siddhartha Pudasaini, Chetan Phakami Pun
arXiv:2507.23248v2 Announce Type: replace-cross
Abstract: Bengali is spoken by more than 230 million people, yet no standardized instrument evaluates large language models (LLMs) on Bengali across th...
By Shimanto Bhowmik, Tawsif Tashwar Dipto, Md Sazzad Islam, Sheryl Hsu, Tahsin Reasat
Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost exclusively in English. As AI agents are globally deployed to a linguistically diverse user base, whether agentic competence measured in English transfers to other languages remains an open question.
arXiv:2608. 11891v1 Announce Type: cross Abstract: Governments increasingly fund indigenous foundation models to strengthen national AI capability, digital sovereignty, and multilingual computing.
By Avinash Agarwal, Vridhi Jain
arXiv:2609.38574v1 Announce Type: new
Abstract: Large language models are frequently proposed as a route to AI-powered services for African communities, but they are least reliable exactly where the...
By Fendji K. E. Jean Louis
The paper examines how four stakeholder groups—AI developers, professional translators, language learners, and language service providers—discuss machine translation on social media. Using a dataset of 79,286 posts from Reddit, Facebook, Bluesky, and Mastodon (2019‑2025), the authors find frequent disagreements and strong conflicts over translation quality, efficiency, and reliability. These conflicts arise because AI communities view the issues as technical, while non‑AI users prioritize quality nuances, time savings, trust, and broader social concerns.
By Yujun Wang, Ehud Reiter, Shimei Pan, Steffen Eger, Wei Zhao
arXiv:2606. 03304v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly evaluated in multilingual settings, yet their inference behavior in low-resource African languages remains underexplored especially under pure prompting without fine-tuning.
By Anuj Tiwari, Terry Oko-odion, Hannah Nwokocha