arXiv:2606. 05176v1 Announce Type: cross Abstract: While large language models (LLMs) show strong performance in natural language understanding and generation, their evaluation and adaptation to domain-specific constraints in telecommunications customer support remain limited.
By Lucas Tamic, Ilan Jaffeux-Cheniout, Xavier Marjou
arXiv:2609.22241v1 Announce Type: new
Abstract: We present H2LooP Telecom Model v1, a domain-specialized large language models fine-tuned for the telecommunications industry. We release two domain-ad...
By Amit Singh, Vedant Nipane, Mayank Goel, Pulkit Agrawal, Sairanjan Mishra
TelecomGPT‑R1‑9B is an open‑source large language model designed specifically for telecom reasoning tasks. It was trained on a 67,427‑example supervised fine‑tuning corpus that covers protocol, knowledge, modeling, and fault reasoning, and further refined with a two‑stage post‑training process involving low‑rank adaptation and policy optimization. The model tops the GSMA open telco leaderboard and matches state‑of‑the‑art closed‑source reasoners across seven public telecom benchmarks.
By Bohao Wang, Chenwei Wu, Haoyu Li, Hang Zou, Yu Tian, Lina Bariah, Li Wei, Chongwen Huang, Yongliang Shen, Zhaoyang Zhang, Merouane Debbah
TelecomGPT‑R1 is an open‑source family of unified telecom reasoning models that address the limitations of existing telecom‑specific and general‑purpose LLMs. It is built on a four‑axis framework—protocol, knowledge, modeling, and fault—and trained on a curated corpus of 104,880 verified question‑answer pairs with chain‑of‑thought reasoning. After supervised fine‑tuning, dynamic sampling policy optimization with task‑routed rubric rewards is used to stabilize reinforcement learning across heterogeneous telecom tasks, achieving an 89.64% mean score on seven GSMA Open Telco Leaderboard benchmarks, surpassing leading proprietary models.
By Bohao Wang, Chenwei Wu, Hang Zou, Yu Tian, Lina Bariah, Li Wei, Chongwen Huang, Yongliang Shen, Zhaoyang Zhang, Merouane Debbah
TeleTables is a benchmark that evaluates large language models on interpreting telecom tables from 3GPP specifications. It contains 2,220 tables in four formats and 500 human‑verified multiple‑choice questions that range from simple retrieval to multi‑step reasoning. Tests on 20 open‑weight LLMs show that closed‑book performance is limited by domain knowledge, while providing the table as context yields high accuracy that still drops with deeper reasoning, evidence scope, and structural complexity.
By Anas Ezzakri, Nicola Piovesan, Mohamed Sana, Antonio De Domenico, Fadhel Ayed, Haozhe Zhang
arXiv:2607. 04071v1 Announce Type: cross Abstract: Portuguese remains underrepresented in text embedding evaluation, despite being one of the most widely spoken languages in the world.
By Lucas Hideki Takeuchi Okamura, Alexandre Alcoforado, Anna Helena Reali Costa
The paper introduces CRAFT, a data‑centric fine‑tuning approach that aligns small language models (SLMs) for pre‑hoc reasoning in AI‑native 6G radio access networks (RAN). By automatically generating verified (input, trace, label) triplets and fine‑tuning with low‑rank adaptation, CRAFT achieves high accuracy and F1 scores on TRACTOR and IC xApp datasets while avoiding parse failures that plague RL methods like GRPO. It also reduces energy consumption by 59% compared to GRPO baselines, offering a more sustainable path to auditable AI in 6G RAN.
By Pranshav Gajjar, Vijay K Shah
arXiv:2606. 13647v1 Announce Type: cross Abstract: We introduce SkMTEB, the first comprehensive MTEB-style text embedding benchmark for Slovak, a low-resource West Slavic language, comprising 31 datasets across 7 task types -- nearly 4$\times$ the depth of existing multilingual benchmark coverage for Slovak.
By Marek \v{S}uppa, Andrej Ridzik, Daniel Hl\'adek, Nat\'alia K\v{n}a\v{z}ekov\'a, Vikt\'oria Ondrejov\'a
FLoKD is an adaptive knowledge‑distillation framework designed for federated fine‑tuning of low‑rank LLMs over wireless networks. It transmits intermediate LoRA activations instead of full parameters or token‑level logits, and uses transformer block importance scoring plus dataset selection to reduce communication. Experiments on WikiText‑103, PTB, and Dialog show a 50‑65% reduction in communication while maintaining competitive perplexity.
By Xinlu Zhang, Na Yan, Yang Su, Yansha Deng, Toktam Mahmoodi
arXiv:2606. 15963v1 Announce Type: cross Abstract: Federated fine-tuning of large language models using parameter-efficient methods such as LoRA enables privacy-preserving adaptation of foundation models.
By Muhammad Waseem, Nurbek Tastan, Andrej Jovanovic, Nicholas D. Lane, Nils Lukas, Karthik Nandakumar, Samuel Horvath
arXiv:2607. 20510v1 Announce Type: new Abstract: We introduce Telco-GAIA, a bilingual, multi-modal benchmark for evaluating tool-using agents on the data of a real-world telecommunications operator.
By Dmitrii Khizbullin, Zaid Alyafeai, Abdelrahman Eldesokey, Nourah AlSultan, Raghad Alshalan, David R. Pugh, Bernard Ghanem
arXiv:2609.23191v1 Announce Type: new
Abstract: Speech-text space alignment is a multimodal representation learning method consisting to map different speech and text into a shared representation spa...
By Yannick Yomie Nzeuhang, Marie Tahon, Paulin Melatagia Yonta