arXiv:2606. 05176v1 Announce Type: cross Abstract: While large language models (LLMs) show strong performance in natural language understanding and generation, their evaluation and adaptation to domain-specific constraints in telecommunications customer support remain limited.
By Lucas Tamic, Ilan Jaffeux-Cheniout, Xavier Marjou
arXiv:2609.22241v1 Announce Type: new
Abstract: We present H2LooP Telecom Model v1, a domain-specialized large language models fine-tuned for the telecommunications industry. We release two domain-ad...
By Amit Singh, Vedant Nipane, Mayank Goel, Pulkit Agrawal, Sairanjan Mishra
TelecomGPT‑R1‑9B is an open‑source large language model designed specifically for telecom reasoning tasks. It was trained on a 67,427‑example supervised fine‑tuning corpus that covers protocol, knowledge, modeling, and fault reasoning, and further refined with a two‑stage post‑training process involving low‑rank adaptation and policy optimization. The model tops the GSMA open telco leaderboard and matches state‑of‑the‑art closed‑source reasoners across seven public telecom benchmarks.
By Bohao Wang, Chenwei Wu, Haoyu Li, Hang Zou, Yu Tian, Lina Bariah, Li Wei, Chongwen Huang, Yongliang Shen, Zhaoyang Zhang, Merouane Debbah
TelecomGPT‑R1 is an open‑source family of unified telecom reasoning models that address the limitations of existing telecom‑specific and general‑purpose LLMs. It is built on a four‑axis framework—protocol, knowledge, modeling, and fault—and trained on a curated corpus of 104,880 verified question‑answer pairs with chain‑of‑thought reasoning. After supervised fine‑tuning, dynamic sampling policy optimization with task‑routed rubric rewards is used to stabilize reinforcement learning across heterogeneous telecom tasks, achieving an 89.64% mean score on seven GSMA Open Telco Leaderboard benchmarks, surpassing leading proprietary models.
By Bohao Wang, Chenwei Wu, Hang Zou, Yu Tian, Lina Bariah, Li Wei, Chongwen Huang, Yongliang Shen, Zhaoyang Zhang, Merouane Debbah
TeleTables is a benchmark that evaluates large language models on interpreting telecom tables from 3GPP specifications. It contains 2,220 tables in four formats and 500 human‑verified multiple‑choice questions that range from simple retrieval to multi‑step reasoning. Tests on 20 open‑weight LLMs show that closed‑book performance is limited by domain knowledge, while providing the table as context yields high accuracy that still drops with deeper reasoning, evidence scope, and structural complexity.
By Anas Ezzakri, Nicola Piovesan, Mohamed Sana, Antonio De Domenico, Fadhel Ayed, Haozhe Zhang
arXiv:2607. 04071v1 Announce Type: cross Abstract: Portuguese remains underrepresented in text embedding evaluation, despite being one of the most widely spoken languages in the world.
By Lucas Hideki Takeuchi Okamura, Alexandre Alcoforado, Anna Helena Reali Costa