arXiv AI

EMS: Multi-Agent Voting via Efficient Majority-then-Stopping

arXiv:2604. 02863v2 Announce Type: replace Abstract: Majority voting is the standard for aggregating multi-agent responses into a final decision.

arXiv AI
2d ago

Revision-Aware Independent Agent Graphs for Dynamic Reasoning

The paper introduces Revision‑Aware Independent Agent Graphs (RIAG) to address dynamic task routing, where an event stream continually revises task bindings and a system must select the correct document version at query time. By repurposing six benchmarks into over 31,000 dynamic episodes, the authors demonstrate that RIAG balances recomputation and reuse, achieving 54.24 % joint routing‑and‑answer accuracy with only 0.62 calls per query—substantially better than the strongest baseline. The study highlights the trade‑off between stale conclusions and wasted work in dynamic reasoning settings.

By Yan Luo, Selim-Antoine Lali, Jeremy Moebel, Iliass Khoutaibi, Ahmadou Aidara, Mengyu Wang
arXiv AI
Aug 18

Agentic Test-Time Scaling for WebAgents

arXiv:2602. 12276v2 Announce Type: replace Abstract: Test-time scaling has become a standard way to improve performance and boost reliability of neural network models.

By Nicholas Lee, Lutfi Eren Erdogan, Chris Joseph John, Surya Krishnapillai, Michael W. Mahoney, Kurt Keutzer, Amir Gholami
arXiv AI
2d ago

Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard

Agent evaluations increasingly benchmark LLMs, but rankings can be swayed by evaluation conditions such as scaffolds or tasks, making reliability claim‑dependent. A Bayesian variance‑decomposition framework applied to 22 benchmarks shows that reliability varies with the measurement goal: fixed model‑scaffold systems rank reliably, while underlying‑model rankings are less stable. Scaffold choice can alter conclusions, and adding more tasks only modestly improves reliability when scaffold coverage is limited; however, pooling diverse benchmarks can substantially raise cross‑task ranking reliability and reduce cost.

By Michael Hardy, Ruhana Azam, Anka Reuel, Mykel Kochenderfer, Sanmi Koyejo
arXiv AI
Sep 4

Auditing Multi-Agent LLM Reasoning Trees Outperforms Majority Vote and LLM-as-Judge

The paper introduces AgentAuditor, a method that improves multi-agent large language model (LLM) reasoning by structuring agent outputs into a Reasoning Tree that captures agreements and divergences, rather than relying on simple majority voting. AgentAuditor resolves conflicts by comparing evidence at key divergence points, enabling efficient localized verification. The authors also propose Anti-Consensus Preference Optimization (ACPO) to train the adjudicator with evidence-verified supervision, reducing reliance on misleading majority cues. Across four MAS frameworks and multiple reasoning benchmarks, AgentAuditor consistently outperforms majority voting, achieving up to 5% absolute accuracy gains while remaining token‑efficient.

By Wei Yang, Shixuan Li, Heng Ping, Peiyu Zhang, Paul Bogdan, Jesse Thomason
arXiv Machine Learning
5d ago

Reinforcement Learning of Communication in a Mesh of Small Language Models

The paper introduces TalkMesh, a decentralized network of small language model agents that learn to communicate effectively during inference. Each agent proposes an answer, scores it with a confidence head, and the most confident agent broadcasts a hint; lower‑confidence agents revise their proposals if a new suggestion scores higher. This gossip‑based consensus, trained via group relative policy optimization, enables a mesh of three agents to match the accuracy of majority voting over 32 samples, and scales to larger meshes to significantly boost performance on benchmarks like GSM8K and MATH-500.

By Mehmet Kerem Turkcan