Hugging Face Trending Papers

MARS: Multi-Specialist LLM Relay System for Competitive Programming

arXiv AI
Aug 26

MARS: Multi-Specialist LLM Relay System for Competitive Programming

MARS (Multi-Agent Relay of Specialized LLMs) is a prompt-only framework that assigns specialized LLM agents—each focused on a particular algorithmic domain such as dynamic programming, graphs, or geometry—to collaboratively solve competitive programming problems. Retrieval-augmented generation selects a small team of relevant specialists for each problem, and the agents iteratively refine a C++17 solution through sandboxed testing, passing structured packets between them until a final infrastructure-fixer normalizes the code. On the CodeContests benchmark, MARS achieves a pass rate of 0.624 with Gemma 4, improving over direct prompting by 14.4 percentage points while reducing wall‑clock cost and token‑spend variance compared to CodeSIM.

By Andrei Mikhailov, Mikhail Burtsev, Alsu Sagirova
arXiv AI
Jul 1

ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents

arXiv:2606. 31174v1 Announce Type: new Abstract: Production large language-model (LLM) agents are increasingly deployed not as lone problem-solvers but as managers: a main model creates specialized subagents, delegates work, and orchestrates their parallel, asynchronous returns through dynamic workflows.

By Kaiwen Xiong, Haonian Ji, Shi Qiu, Zeyu Zheng, Cihang Xie, Xinyu Ye, Huaxiu Yao
arXiv AI
Jul 24

DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP Servers

arXiv:2607. 20531v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed over Model Context Protocol (MCP) servers, yet the benchmarks used to evaluate them score the final answer or a fixed "ground-truth" list of tools, both of which are fragile once the underlying data is live and stateful.

By Jerzy Kami\'nski, Ilya Galyukshev, Artem Kuznetsov, Sergey Chuprin, Kirill Redko, Aidar Shumbalov, Anna Kalyuzhnaya
Hugging Face Trending Papers
Aug 6

CodeGrep: An RL-Trained Retrieval Agent for LLM Coding Agents

Modern LLM coding agents such as Claude Code and OpenHands share a common inefficiency: they spend much of their token budget finding the file to patch, rather than patching it. On SWE-Bench Verified, a 30B OpenHands agent averages 23 rounds and 631K tokens per resolved issue, with many calls spent on grep, glob, and view_file during repository exploration.

arXiv AI
3d ago

Zero2Repo: Can Coding Agents Build Repositories from Scratch?

arXiv:2609.38269v1 Announce Type: cross Abstract: Coding agents are increasingly asked to build software rather than patch it, yet benchmarks for from-scratch repository construction are mostly limit...

By Pei Yang, Tianyu Shi, Yuhang Yao, Wanyi Chen, Tongyun Yang, Dun Pei, Haonan Wang, Pengbin Feng, Guanxu Yu, Jingchun Huang, Zeyu Zhang, Shuhan Sun, Hao Li, Xiang Li, Jie Xiao, Xinyu Wang, Hanxin Chen, Daqi Li, Qi Jia, Hongshan Lin, Zhizhou Gu, Zijun Tian, Weizhi Du, Lynn Ai, Eric Yang