arXiv:2608. 10692v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions.
By Junjie Ye, Zhuohui Sheng, Shaofan Liu, Yulun Zhu, Wenjie Fu, Dingwei Zhu, Ming Zhang, Yujiong Shen, Weichao Wang, Xin Zhao, Shihan Dou, Tao Gui, Qi Zhang, Xuanjing Huang, Pluto Zhou
arXiv:2608.23078v1 Announce Type: new
Abstract: Large language models increasingly operate over large collections of tools, functions, APIs, and specialized agents. As the candidate action space grow...
By Saurav Singla, Aarav Singla, Advik Gupta, Parnika Gupta
The paper introduces CRAFT, a data‑centric fine‑tuning approach that aligns small language models (SLMs) for pre‑hoc reasoning in AI‑native 6G radio access networks (RAN). By automatically generating verified (input, trace, label) triplets and fine‑tuning with low‑rank adaptation, CRAFT achieves high accuracy and F1 scores on TRACTOR and IC xApp datasets while avoiding parse failures that plague RL methods like GRPO. It also reduces energy consumption by 59% compared to GRPO baselines, offering a more sustainable path to auditable AI in 6G RAN.
By Pranshav Gajjar, Vijay K Shah
The next generation of mobile networks is envisioned as fully AI-native, with AI-RAN architectures embedding small language models (SLMs) to perform reasoning over real-time telemetry. The state-of-th...
arXiv:2608.22472v1 Announce Type: new
Abstract: Function calling represents the core capability of agentic large language models (LLMs). Existing research has focused on enhancing LLMs function-calli...
By Yalda Taheri, Mohammad Hassan Heydari, Erfan Naaman, Afsaneh Fatemi
Agent Seer is a pipeline that automatically synthesizes realistic evaluation scenarios for AI agents that use external tools, using only the tool’s specification (function names, natural‑language descriptions, and typed parameter schemas). Starting from a single Model Context Protocol (MCP) specification, it enriches raw schemas, generates graded scenarios with synthetic tool outputs, and expands them into mock‑data‑grounded multi‑turn dialogues that demonstrate strong tool‑calling correctness and conversational coherence. Across seven diverse MCP specifications, the pipeline achieves high quality, with parameter‑schema complexity emerging as the main driver of quality variation and argument‑value accuracy identified as the dominant failure mode.
By Harish Karumuri, Mahesh Vemula, David Lopes Pegna
The paper presents a unified evaluation of seven open reasoning language models across four benchmarks (ARC-Challenge, GSM8K, MATH levels 1–3, and TruthfulQA MC1) using a consistent 238-example subset and three prompting strategies (zero-shot, chain-of-thought, few-shot CoT). It reports not only accuracy but also Wilson confidence intervals, latency, VRAM usage, weighted aggregate performance, Pareto-efficient points, prompt-sensitivity, and compatibility diagnostics, revealing that Gemma-4-26B-A4B tops the weighted score while Gemma-4-E4B offers a strong practical trade-off. The study emphasizes that model rankings shift with prompting strategy and that deployment trade-offs remain crucial, advocating for a deployment-aware, multi-objective evaluation framework rather than a single-score leaderboard.
By Md Motaleb Hossen Manik, Ge Wang
arXiv:2601. 22588v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are widely used as reference-free evaluators via prompting, but this "LLM-as-a-Judge" paradigm is costly, opaque, and sensitive to prompt design.
By Zhuochun Li, Yong Zhang, Ming Li, Yuelyu Ji, Yiming Zeng, Ning Cheng, Yun Zhu, Yanmeng Wang, Shaojun Wang, Jing Xiao, Daqing He
arXiv:2606. 25156v3 Announce Type: replace-cross Abstract: Native length extrapolation remain a weakly solvable problem in language modeling due to trade-off balancing between exact retrieval fidelity, long-document likelihood, and inference efficiency.
By Habibullah Akbar
arXiv:2606. 12387v1 Announce Type: cross Abstract: Large Language Models (LLMs) have democratized database access through Text-to-SQL, but moving from prototypes to production remains difficult.
By Zhiyi Chen, Jie Song, Peng Li
arXiv:2603. 29025v3 Announce Type: replace-cross Abstract: Large language models fail when a salient surface cue conflicts with an unstated feasibility constraint.
By Yubo Li, Lu Zhang, Tianchong Jiang, Ramayya Krishnan, Rema Padman
arXiv:2606. 12117v1 Announce Type: cross Abstract: Benchmark scores often misrepresent a large language model's (LLM's) knowledge, because they rely, e.
By Selen Erkan, Bastian Boll, Kristian Kersting, Bj\"orn Deiseroth, Letitia Parcalabescu