Hugging Face Trending Papers

TRACE: Tackling Real-World Resource Assignment Problems via Agentic Heuristic Design

arXiv Machine Learning
1d ago

TRACE: Tackling Real-World Resource Assignment Problems via Agentic Heuristic Design

TRACE tackles real‑world dynamic resource assignment by combining evolutionary automatic heuristic design with an agentic knowledge‑extraction workflow. A Reasoner agent interprets system logs to hypothesize about underlying dynamics, while a Coder agent generates and runs schema‑specific code to validate these hypotheses, producing insights or executable tools for the evolved heuristics. Evaluations on a synthetic cloud benchmark and a 5G vRAN scenario show that TRACE outperforms existing AHD methods, delivering more auditable heuristics with less than 2% overhead.

By Jose A. Ayala-Romero, Andres Garcia-Saavedra, Xavier Costa-Perez
arXiv AI
Sep 4

Evolving Excellence: Automated Optimization of LLM-based Agents

The paper introduces ARTEMIS, a no-code evolutionary optimization platform that automatically tunes large language model (LLM) agents by jointly optimizing prompts, tool descriptions, and parameters using semantically-aware genetic operators. Starting from a benchmark script and natural language goals, ARTEMIS discovers configurable components, extracts performance signals from execution logs, and evolves configurations without architectural changes. Experiments on four agent systems show significant gains: a 13.6% increase in acceptance rate for the ALE Agent, a 10.1% performance boost for the Mini‑SWE Agent, a 36.9% token‑reduction for the CrewAI Agent, and a 22% accuracy improvement for the MathTales‑Teacher Agent using a smaller open‑source model.

By Paul Brookes, Vardan Voskanyan, Rafail Giavrimis, Matthew Truscott, Mina Ilieva, Chrystalla Pavlou, Alexandru Staicu, Manal Adham, Will Evers- Hood, Jingzhi Gong, Kejia Zhang, Matvey Fedoseev, Vishal Sharma, Roman Bauer, Zheng Wang, Hema Nair, Wei Jie, Tianhua Xu, Aurora Constantin, Leslie Kanthan, Michail Basios
arXiv AI
Sep 18

AgentPProf: Semantic Profiler for Long Horizon AI Agents

AgentPProf is a new semantic profiler designed for long‑horizon AI agents that aggregates agent trajectories into pprof‑compatible profiles, enabling flame‑graph visualization and hierarchical attribution of tasks and subtasks. It introduces a semantic operation stack model and recursive operation segmentation to replace traditional call‑stack profiling, addressing the challenge of profiling agent intent rather than code paths. In evaluations, AgentPProf achieves high F1 scores against human annotations and significantly improves problem‑localization metrics, demonstrating its effectiveness in attributing resources, locating issues, and optimizing token cost.

By Yusheng Zheng, Chaokun Chang, Yu Mao, Tianyuan Wu, Yuxi Huang, Tao Ma, Wenan Mao, Shuyi Cheng, Andi Quinn, Wei Wang
arXiv AI
Aug 26

PeakBench: Benchmarking Resource-Aware Tool Invocation in LLM Agents

PeakBench is a new benchmark designed to evaluate how large language model agents invoke multiple tools while respecting resource constraints and parallel execution. It provides executable multi‑tool workflows with dependency annotations and measured resource profiles, and introduces a two‑part evaluation framework that separates logical planning from physical scheduling. The study shows that strong logical planning alone does not guarantee safe or efficient execution, and that providing resource information can reduce overflows and improve utilization.

By Zhi-Kai Chen, Xu-Xiang Zhong, Song-Yan Li, De-Chuan Zhan, Han-Jia Ye
arXiv AI
Sep 2

UniACE: A Unified Framework for Evaluating LLM Agentic Capabilities

UniACE is a unified framework that standardizes the evaluation of large language model (LLM) agents by representing each benchmark as an instruction–tool–environment triplet and running models through a shared, task‑agnostic harness in isolated runtimes. It preserves native success criteria, offers an offline mode for dynamic‑resource tasks, and standardizes efficiency metrics, execution records, and failure attribution. Applying UniACE to 7 benchmarks across 24 domains and 15 models revealed significant score shifts, ranking reversals, and sensitivity to evidence representation, highlighting the impact of evaluation configuration on reported agent performance.

By Pengyu Zhu, Lijun Li, Yaxing Lyu, Qianxin Luo, Jingyi Yang, Yi Liu, Tingfeng Hui, Xinyu Yuan, Li Sun, Sen Su, Jing Shao
arXiv AI
Aug 26

Beyond Executable Models: The Pufibara Agent Harness and the Modelica Agent Workflow Benchmark for Physical System Modeling

The paper introduces Pufibara, an agent harness designed to maintain engineering state and evidence across revisions in Modelica-based physical system modeling. It also presents a 232-task Modelica Agent Workflow Benchmark covering model repair, generation, and tuning, evaluated by an external benchmark-owned evaluator. Experiments show Pufibara outperforms Claude Code in task success and resource efficiency across two LLM backends.

By Zizhe Wang
arXiv AI
Aug 18

From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems

arXiv:2608. 15127v1 Announce Type: cross Abstract: Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, environments, and persistent state.

By Chaokun Chang, Yukun Zhou, Kaihua Fu, Dakai An, Tianyu Feng, Hanfeng Lu, Sheng Yao, Pu Guo, Yinghao Yu, Yizhou Shan, Bo Li, Binhang Yuan, Wei Wang