arXiv AI
Sep 2

UniACE: A Unified Framework for Evaluating LLM Agentic Capabilities

UniACE is a unified framework that standardizes the evaluation of large language model (LLM) agents by representing each benchmark as an instruction–tool–environment triplet and running models through a shared, task‑agnostic harness in isolated runtimes. It preserves native success criteria, offers an offline mode for dynamic‑resource tasks, and standardizes efficiency metrics, execution records, and failure attribution. Applying UniACE to 7 benchmarks across 24 domains and 15 models revealed significant score shifts, ranking reversals, and sensitivity to evidence representation, highlighting the impact of evaluation configuration on reported agent performance.

By Pengyu Zhu, Lijun Li, Yaxing Lyu, Qianxin Luo, Jingyi Yang, Yi Liu, Tingfeng Hui, Xinyu Yuan, Li Sun, Sen Su, Jing Shao
arXiv Machine Learning
Sep 14

HoliBench: A Cross-Platform Benchmarking and Deployment Toolkit for Foundation Models in CPS-IoT Applications

HoliBench is a modular benchmarking and deployment toolkit that jointly measures accuracy, latency, and energy for foundation models across a wide range of devices, from single-board computers to GPU servers. It provides a platform abstraction layer that calibrates cross-device measurements and supports multiple model modalities, inference engines, and quantization levels. Using HoliBench, the authors evaluated 20 models on 7 device types, revealing tradeoffs such as limited latency gains from quantization on low‑precision hardware and diminishing accuracy returns relative to energy consumption, while also showing that single-model profiles can predict multi-model pipeline performance within a few percent.

By Inesh Chakrabarti, Zejun Xiong, Pragya Sharma, Mani Srivastava
arXiv AI
Aug 19

Beyond FLOPs: Energy-Aware Knowledge Distillation for Sustainable LLMs on Code-Related Task

The paper explores energy-aware knowledge distillation for large language models (LLMs) used in software engineering tasks such as clone detection, vulnerability prediction, and code summarization. It shows that the commonly used FLOPs metric does not reliably reflect actual energy consumption, and that using energy-surrogate models during distillation can reduce inference energy by up to 90% and memory usage by 86% with only modest accuracy loss. The study demonstrates that guiding distillation with direct energy estimates improves the sustainability and deployability of LLMs on consumer hardware.

By Enrique Barba Roque, Lu\'is Cruz, Annibale Panichella