The paper proposes that future 6G networks will use Large Language Model agents to manage the Radio Access Network, but current designs mistakenly treat inter‑agent messages as objective facts. It argues that messages are actually traces of the sender’s reasoning, carrying subjective conclusions that can propagate hallucinations and cause outages. By modeling these interactions as cognitive channels on a cellular sheaf, the authors derive five design principles—treating messages as evidence of hidden reasoning, defining trust as a continuous cognitive Signal‑to‑Noise Ratio, computing network consistency via the sheaf’s Laplacian, limiting peer‑modeling to two levels, and bounding credible capacity by goal alignment—and validate them with a signaling‑storm study on 1B‑parameter telecom language models.
By Hatim Chergui, Carolina Fern\'{a}ndez-Mart\'{i}nez, Mehdi Bennis, Merouane Debbah
arXiv:2606. 03034v1 Announce Type: cross Abstract: Large language model (LLM) agents have begun to delegate work to one another.
By Gaurav Naresh Mittal
arXiv:2607. 03598v1 Announce Type: cross Abstract: When a person shares something with a language model, the model often answers the surface of the message rather than what the sender was doing by sending it: share a finished project and it critiques the code; share a raw late-night line and it runs a wellness check.
By Alex Kwon
The study investigates how limited reading capacity and claim wording influence consensus outcomes in language‑model networks. By modeling message capacity as the number of messages an agent reads, the authors show that when agents read fewer than about 6.4 messages on average, a wrong consensus becomes unreachable. However, the wording of a claim—its inherent threshold—can override this effect, leading to incorrect consensus even when most agents start correct.
By Makoto Fukushima
The paper introduces TalkMesh, a decentralized network of small language model agents that learn to communicate effectively during inference. Each agent proposes an answer, scores it with a confidence head, and the most confident agent broadcasts a hint; lower‑confidence agents revise their proposals if a new suggestion scores higher. This gossip‑based consensus, trained via group relative policy optimization, enables a mesh of three agents to match the accuracy of majority voting over 32 samples, and scales to larger meshes to significantly boost performance on benchmarks like GSM8K and MATH-500.
By Mehmet Kerem Turkcan
arXiv:2606. 02646v1 Announce Type: cross Abstract: Inference-time multi-agent LLM scaling lacks a shared unit: counting nominal agents conflates cost with independent evidence.
By Bla\v{z} Bertalani\v{c}, Carolina Fortuna
arXiv:2604.27167v3 Announce Type: replace-cross
Abstract: On the named Prisoner's Dilemma under direct prompting, three larger instruction-tuned models, Llama-3-70B, Qwen2.5-32B, and Qwen2.5-72B, loc...
By Paraskevas V. Lekeas, Giorgos Stamatopoulos
arXiv:2607. 16133v1 Announce Type: cross Abstract: LLM powered multi-agent systems (MAS) have emerged as a promising paradigm for complex tasks.
By Wendi Yu, Lianhao Zhou, Xiangjue Dong, Sai Sudarshan Barath, Declan Staunton, Byung-Jun Yoon, Xiaoning Qian, James Caverlee, Shuiwang Ji
arXiv:2609.34373v2 Announce Type: replace
Abstract: Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message s...
By Mihir Chauhan, Aniket Bera
arXiv:2606. 14476v1 Announce Type: new Abstract: A growing line of work equips large language model (LLM) agents with graph neural networks (GNNs) as callable tools, assuming the agent exercises judgment over when and how much to rely on such a tool.
By Zhongyuan Wang, Pratyusha Vemuri
arXiv:2606. 27409v1 Announce Type: cross Abstract: Multi-agent large language model (LLM) systems often rely on verifier and critic agents to suppress hallucinations, but verification is delayed.
By Igor Itkin
The paper introduces a Bayesian self‑escalation strategy for hierarchical large‑language‑model agents, allowing an agent to detect during its own reasoning that it is unlikely to succeed and hand control over to a stronger model. The authors formalise this as an optimal‑stopping problem over a learned competence posterior, derive a myopic escalation threshold, and prove that the optimal policy is a time‑varying threshold without assumptions on the raw signal. They provide theoretical guarantees—including a 1/√n regret decay with n calibration trajectories—and validate the approach in simulations and a real‑model code‑generation cascade, showing that the escalation frontier outperforms post‑hoc routing at equal cost.
whyItMatters":"The study offers a principled, theoretically grounded method for agents to dynamically decide when to seek stronger models, potentially improving efficiency and reliability in hierarchical LLM systems."
By Nadeem Shaikh