Bringing Large Language Models (LLMs) into industrial ride-hailing dispatch as semantic feature extractors over platform-scale behavioral logs is a compelling but under-explored data systems problem. Production matching pipelines remain dominated by structured numerical features, yet decisive behavioral signals (e.
AgentPProf is a new semantic profiler designed for long‑horizon AI agents that aggregates agent trajectories into pprof‑compatible profiles, enabling flame‑graph visualization and hierarchical attribution of tasks and subtasks. It introduces a semantic operation stack model and recursive operation segmentation to replace traditional call‑stack profiling, addressing the challenge of profiling agent intent rather than code paths. In evaluations, AgentPProf achieves high F1 scores against human annotations and significantly improves problem‑localization metrics, demonstrating its effectiveness in attributing resources, locating issues, and optimizing token cost.
By Yusheng Zheng, Chaokun Chang, Yu Mao, Tianyuan Wu, Yuxi Huang, Tao Ma, Wenan Mao, Shuyi Cheng, Andi Quinn, Wei Wang
TRACE tackles real‑world dynamic resource assignment by combining evolutionary automatic heuristic design with an agentic knowledge‑extraction workflow. A Reasoner agent interprets system logs to hypothesize about underlying dynamics, while a Coder agent generates and runs schema‑specific code to validate these hypotheses, producing insights or executable tools for the evolved heuristics. Evaluations on a synthetic cloud benchmark and a 5G vRAN scenario show that TRACE outperforms existing AHD methods, delivering more auditable heuristics with less than 2% overhead.
By Jose A. Ayala-Romero, Andres Garcia-Saavedra, Xavier Costa-Perez
Dynamic resource assignment, the real-time allocation of task streams to heterogeneous processing nodes, is the backbone of modern computing infrastructure. While learning-based schedulers excel in re...
arXiv:2608. 11241v1 Announce Type: new Abstract: Deploying LLM agents into industrial recommender operations exposes a three-way tension we frame as the autonomy-determinism-efficiency trilemma: general autonomy (interpreting operator intent, generating glue code zero-shot), industrial determinism (schema-conforming feature extraction, non-crashing A/B, zero compliance-path hallucination), and end-to-end efficiency.
By Dongyang Ao, Kaixiang Fang, Shijie Xu
arXiv:2608. 19751v1 Announce Type: new Abstract: Micro-View Order-Dispatching assigns available drivers to passenger orders within each dispatch batch and is critical to the service quality and operational efficiency of ride-hailing platforms.
By Chuang Liu, Yuxueqing Zhang, Tengfei Lyu, Zirui Yuan, Weiqi Hu, Yanghan Cheng, Ming Wang, Li Ma, Zihao Lu
UniACE is a unified framework that standardizes the evaluation of large language model (LLM) agents by representing each benchmark as an instruction–tool–environment triplet and running models through a shared, task‑agnostic harness in isolated runtimes. It preserves native success criteria, offers an offline mode for dynamic‑resource tasks, and standardizes efficiency metrics, execution records, and failure attribution. Applying UniACE to 7 benchmarks across 24 domains and 15 models revealed significant score shifts, ranking reversals, and sensitivity to evidence representation, highlighting the impact of evaluation configuration on reported agent performance.
By Pengyu Zhu, Lijun Li, Yaxing Lyu, Qianxin Luo, Jingyi Yang, Yi Liu, Tingfeng Hui, Xinyu Yuan, Li Sun, Sen Su, Jing Shao
The paper introduces KOPE, an experience‑driven framework that records hardware kernel optimization trajectories in an Experience Graph Memory and uses Active Context Management and Injection to retrieve relevant past decisions under a fixed token budget. KOPE preserves decision order, outcomes, and alternative branches, enabling evidence from completed runs to inform future optimization steps. In experiments, KOPE achieves a 1.54× speedup over the strongest baseline, raises pass rates from 60.0% to 84.6%, and reduces token consumption dramatically, demonstrating the benefits of continual learning from external experience while keeping the foundation model unchanged.
By Siyuan Chen, Runlin Hou, Shenxiu Wu, Yansong Sun, Junming Cao, Yiyu Zhang, Shudi Shao, Junhao Qiu, Zhichao Lu, Qingfu Zhang
arXiv:2603. 03589v3 Announce Type: replace-cross Abstract: Recent advances in large language models (LLMs) transform how machine learning (ML) pipelines are developed and evaluated.
By Arnab Phani, Elias Strauss, Sebastian Schelter
AgentPerfBench is a new benchmarking suite designed to evaluate the inference performance of agentic large language models (LLMs) that handle multi‑turn, tool‑using, and context‑expanding tasks. It builds on real traces from agentic benchmarks such as SWE‑Bench and TerminalBench, and generates synthetic profiles that reflect realistic input/output lengths and turn counts. The suite also provides kernel‑level Nsight Compute traces and a multi‑dimensional roofline model to identify hardware bottlenecks and quantify the gap between traditional chat benchmarks and agentic workloads.
arXiv:2607. 19349v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as always-on online services, making efficient LLM serving a critical systems challenge.
By Tiancheng Zhang, Shaoyuan Huang, Mingyuan Wang, Yunfeng Zhao, Xiaofei Wang, Wenyu Wang
PlannerForge is a unified LLM‑agent framework that covers the entire scenario‑based testing pipeline for autonomous driving systems, from scenario generation to ADS assessment, and adds ADS enhancement and benchmarking stages. It was evaluated with ten off‑the‑shelf LLMs across all tasks and five prompt conditions, achieving best‑per‑task scores between 0.88 and 1.00 and matching commercial APIs with open‑source models such as Qwen3.6:35B. The end‑to‑end chaining retains 83% of seed queries for commercial backends and 78% for open‑source, outperforming existing tools like Scenario Factory 2.0 and BM25 in natural‑language generation, attribute realization, and physically valid edits.
whyItMatters":"PlannerForge demonstrates that a single LLM‑based system can streamline and improve the fragmented scenario‑based testing workflow for autonomous driving, achieving high performance without domain‑specific fine‑tuning."
By Yuan Gao, Sebastian M\"uller, Mattia Piccinini, Marc Kaufeld, Yuchen Zhang, Finn Rasmus Sch\"afer, Qunying Song, Johannes Betz