Hugging Face Trending Papers

ProfiLLM: Utility-Aligned Agentic User Profiling for Industrial Ride-Hailing Dispatch

Bringing Large Language Models (LLMs) into industrial ride-hailing dispatch as semantic feature extractors over platform-scale behavioral logs is a compelling but under-explored data systems problem. Production matching pipelines remain dominated by structured numerical features, yet decisive behavioral signals (e.

arXiv AI
Sep 18

AgentPProf: Semantic Profiler for Long Horizon AI Agents

AgentPProf is a new semantic profiler designed for long‑horizon AI agents that aggregates agent trajectories into pprof‑compatible profiles, enabling flame‑graph visualization and hierarchical attribution of tasks and subtasks. It introduces a semantic operation stack model and recursive operation segmentation to replace traditional call‑stack profiling, addressing the challenge of profiling agent intent rather than code paths. In evaluations, AgentPProf achieves high F1 scores against human annotations and significantly improves problem‑localization metrics, demonstrating its effectiveness in attributing resources, locating issues, and optimizing token cost.

By Yusheng Zheng, Chaokun Chang, Yu Mao, Tianyuan Wu, Yuxi Huang, Tao Ma, Wenan Mao, Shuyi Cheng, Andi Quinn, Wei Wang
arXiv AI
Aug 13

RecSys Factory: Bounding LLM Agent Autonomy to Decision Points in the Industrial Recommender Lifecycle

arXiv:2608. 11241v1 Announce Type: new Abstract: Deploying LLM agents into industrial recommender operations exposes a three-way tension we frame as the autonomy-determinism-efficiency trilemma: general autonomy (interpreting operator intent, generating glue code zero-shot), industrial determinism (schema-conforming feature extraction, non-crashing A/B, zero compliance-path hallucination), and end-to-end efficiency.

By Dongyang Ao, Kaixiang Fang, Shijie Xu
arXiv AI
Sep 2

UniACE: A Unified Framework for Evaluating LLM Agentic Capabilities

UniACE is a unified framework that standardizes the evaluation of large language model (LLM) agents by representing each benchmark as an instruction–tool–environment triplet and running models through a shared, task‑agnostic harness in isolated runtimes. It preserves native success criteria, offers an offline mode for dynamic‑resource tasks, and standardizes efficiency metrics, execution records, and failure attribution. Applying UniACE to 7 benchmarks across 24 domains and 15 models revealed significant score shifts, ranking reversals, and sensitivity to evidence representation, highlighting the impact of evaluation configuration on reported agent performance.

By Pengyu Zhu, Lijun Li, Yaxing Lyu, Qianxin Luo, Jingyi Yang, Yi Liu, Tingfeng Hui, Xinyu Yuan, Li Sun, Sen Su, Jing Shao
arXiv Machine Learning
1d ago

TRACE: Tackling Real-World Resource Assignment Problems via Agentic Heuristic Design

TRACE tackles real‑world dynamic resource assignment by combining evolutionary automatic heuristic design with an agentic knowledge‑extraction workflow. A Reasoner agent interprets system logs to hypothesize about underlying dynamics, while a Coder agent generates and runs schema‑specific code to validate these hypotheses, producing insights or executable tools for the evolved heuristics. Evaluations on a synthetic cloud benchmark and a 5G vRAN scenario show that TRACE outperforms existing AHD methods, delivering more auditable heuristics with less than 2% overhead.

By Jose A. Ayala-Romero, Andres Garcia-Saavedra, Xavier Costa-Perez
arXiv Machine Learning
Aug 27

Beyond Scaling: Self-Evolving LLM Agents for Hardware Kernel Optimization via an Experience-Driven Workflow and Experience Graph Memory

The paper introduces KOPE, an experience‑driven framework that records hardware kernel optimization trajectories in an Experience Graph Memory and uses Active Context Management and Injection to retrieve relevant past decisions under a fixed token budget. KOPE preserves decision order, outcomes, and alternative branches, enabling evidence from completed runs to inform future optimization steps. In experiments, KOPE achieves a 1.54× speedup over the strongest baseline, raises pass rates from 60.0% to 84.6%, and reduces token consumption dramatically, demonstrating the benefits of continual learning from external experience while keeping the foundation model unchanged.

By Siyuan Chen, Runlin Hou, Shenxiu Wu, Yansong Sun, Junming Cao, Yiyu Zhang, Shudi Shao, Junhao Qiu, Zhichao Lu, Qingfu Zhang
arXiv Machine Learning
Jun 2

UME: A Unified Meta-Generalization Framework for Cross-Domain ETA

arXiv:2606. 00979v1 Announce Type: new Abstract: Accurate Estimated Time of Arrival (ETA) prediction on checkout page is crucial in instant logistics for enhancing user satisfaction, optimizing dispatching, and controlling operational costs.

By Duo Wang, Qiong Wu, Jianguo Wu, Ruiyu Xu, Jinhui Yi, Zhonggen Sun, Zhentao Zhang, Yu Zhang, Ke Xing, Yongjun Yin, Zishuo Li, Jianwen Huang
arXiv AI
Sep 7

AutoLR: Automating the Path from Research to Launch Review in Industrial Recommender Systems

AutoLR is an autonomous harness designed to streamline the iterative research‑and‑engineering cycle for industrial recommender systems, exemplified by NetEase’s gaming‑community app DASHEN. It integrates a multi‑expert council for adversarial review, a deterministic evidence‑weighted selector to allocate trial budgets, and a layered knowledge system that fuses external research with domain‑specific insights and empirical evidence. Large language model agents handle semantic reasoning and code generation, while deterministic controllers maintain control over execution, metrics, guardrails, and state management.

By Qi Zhang, Yanlin Chen, Wenchao Xiao