arXiv Computation and Language
Aug 28

TelecomGPT-R1: A Unified Open-Source Reasoner for the Telecom Stack

TelecomGPT‑R1‑9B is an open‑source large language model designed specifically for telecom reasoning tasks. It was trained on a 67,427‑example supervised fine‑tuning corpus that covers protocol, knowledge, modeling, and fault reasoning, and further refined with a two‑stage post‑training process involving low‑rank adaptation and policy optimization. The model tops the GSMA open telco leaderboard and matches state‑of‑the‑art closed‑source reasoners across seven public telecom benchmarks.

By Bohao Wang, Chenwei Wu, Haoyu Li, Hang Zou, Yu Tian, Lina Bariah, Li Wei, Chongwen Huang, Yongliang Shen, Zhaoyang Zhang, Merouane Debbah
arXiv Computation and Language
Sep 23

TelecomGPT-R1: Unified Post-Training for Reasoning Across Heterogeneous Telecom Tasks

TelecomGPT‑R1 is an open‑source family of unified telecom reasoning models that address the limitations of existing telecom‑specific and general‑purpose LLMs. It is built on a four‑axis framework—protocol, knowledge, modeling, and fault—and trained on a curated corpus of 104,880 verified question‑answer pairs with chain‑of‑thought reasoning. After supervised fine‑tuning, dynamic sampling policy optimization with task‑routed rubric rewards is used to stabilize reinforcement learning across heterogeneous telecom tasks, achieving an 89.64% mean score on seven GSMA Open Telco Leaderboard benchmarks, surpassing leading proprietary models.

By Bohao Wang, Chenwei Wu, Hang Zou, Yu Tian, Lina Bariah, Li Wei, Chongwen Huang, Yongliang Shen, Zhaoyang Zhang, Merouane Debbah
arXiv AI
Sep 7

TeleTables: A Benchmark for Large Language Models in Telecom Table Interpretation

TeleTables is a benchmark that evaluates large language models on interpreting telecom tables from 3GPP specifications. It contains 2,220 tables in four formats and 500 human‑verified multiple‑choice questions that range from simple retrieval to multi‑step reasoning. Tests on 20 open‑weight LLMs show that closed‑book performance is limited by domain knowledge, while providing the table as context yields high accuracy that still drops with deeper reasoning, evidence scope, and structural complexity.

By Anas Ezzakri, Nicola Piovesan, Mohamed Sana, Antonio De Domenico, Fadhel Ayed, Haozhe Zhang
arXiv AI
Aug 24

Free-Text Evaluation of LLMs for 5G Domain Knowledge and Fault Analysis using LLM-as-Judge

The paper evaluates lightweight, edge‑deployable large language models—Claude‑Haiku‑4.5, GPT‑5.4‑Mini, and Gemini‑3.1‑Flash‑Lite—on free‑text 5G domain knowledge and fault‑analysis tasks using three benchmarks (TeleQNA ORAN FT, 5G‑Faults FT, TeleInter FT). All models achieve at least 90% accuracy on fault diagnosis, but zero‑shot recall of 3GPP and O‑RAN specifications remains below 60%. Multi‑judge scoring yields a mean inter‑judge agreement of at least 0.90, and Gemini‑3.1‑Flash‑Lite emerges as the most efficient model for production telecom deployments.

By Rishiraj Sengupta, Sotiris Chatzimiltis, Mohammad Shojafar, Xiatian Zhu
arXiv Machine Learning
Sep 2

CRAFT: Fine-Tuning Pre-hoc Explainability in AI-native 6G RAN

The paper introduces CRAFT, a data‑centric fine‑tuning approach that aligns small language models (SLMs) for pre‑hoc reasoning in AI‑native 6G radio access networks (RAN). By automatically generating verified (input, trace, label) triplets and fine‑tuning with low‑rank adaptation, CRAFT achieves high accuracy and F1 scores on TRACTOR and IC xApp datasets while avoiding parse failures that plague RL methods like GRPO. It also reduces energy consumption by 59% compared to GRPO baselines, offering a more sustainable path to auditable AI in 6G RAN.

By Pranshav Gajjar, Vijay K Shah