arXiv AI By Nivasini Ananthakrishnan, Meena Jagadeesan

Power and Limitations of Aggregation in Compound AI Systems

Read the original on arXiv AI →

arXiv:2602. 21556v2 Announce Type: replace Abstract: When designing compound AI systems, a common approach is to query multiple copies of the same model and aggregate the responses to produce a synthesized output.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 28

From Atomic to Agentic: Towards Interpretable Evaluation of LLMs' Agentic Mathematical Capabilities

The paper introduces a new benchmark that evaluates large language models (LLMs) on their agentic mathematical reasoning rather than just final answers. It aligns problem‑solving behaviors with a taxonomy of reusable mathematical atomic capabilities and includes planning, action, and feedback tasks in both textual and multimodal settings. Experiments show that models with similar end‑to‑end accuracy can have very different agentic profiles, highlighting the importance of process‑level evaluation.

By Jiayi Kuang, Yinghui Li, Yunze Song, Keyu Chen, Zhifeng Shen, Yangning Li, Yidong Wang, Di Yin, Ruizhi Qiao, Xing Sun, Kai Jin, Ying Shen, Liang Lin, Philip S. Yu
arXiv AI
Jun 12

From Digital to Physical: Digital Agents as Autonomous Coaches for Physical Intelligence

arXiv:2601. 21570v2 Announce Type: replace Abstract: The field of Embodied AI is witnessing a rapid evolution toward general-purpose robotic systems, fueled by high-fidelity simulation and large-scale data collection.

By Zixing Lei, Genjia Liu, Yuanshuo Zhang, Qipeng Liu, Yuzhu Cai, Sixiang Chen, Jixian Wu, Yunhong Wang, Weixin Li, Chuan Wen, Bo Zhao, Shanghang Zhang, Wenzhao Lian, Siheng Chen
arXiv AI
Sep 15

The Universe of Universes: Benefit Yield Functions, Implosion Thresholds, and Infrastructure-Aware Optimization in Multi-LLM Systems

The paper introduces the Universe of Universes (UoU) framework, treating the ecosystem of major large language models as a structured retrieval corpus and proposing a compositional Automated Reasoning and Machine Learning architecture for cross-model retrieval‑augmented generation. It formally defines the Benefit Yield Function (BYF), measuring marginal performance gain per added model, and identifies an implosion threshold θ* where BYF becomes zero and ensemble performance degrades. The work highlights gaps in current LLM ensemble research, such as lack of performance analysis across full model universes, and connects these findings to implications for DoD AI acquisition policy and testing of AI‑enabled systems.

By Danielle Franklin, Vasu Raj Jain