arXiv Machine Learning

Who Belongs in the Eval Set? A Capability-Taxonomy-Driven Pipeline for Curating Regression Eval Sets in Agent-Extensibility Platforms

arXiv:2608. 01004v1 Announce Type: new Abstract: Platform teams hosting agent-extensibility surfaces face a regression-economics paradox: every onboarding customer ships an evaluation set tuned to their domain, but the platform's regression set must live under a hard query-count ceiling bounded by release cadence.

arXiv AI
Sep 18

Semantic Feature Analysis: Improving Agents Without Searching Over Rollouts

Semantic Feature Analysis (SFA) is a method that refines agent specifications without performing any rollout-based search. It analyzes existing execution traces, clusters workflow node outputs, extracts semantic feature classes via an extended subject‑verb‑object schema, ranks these features with a decision tree, and injects the most impactful features back into the system prompt. Evaluations on four benchmarks show that SFA consistently outperforms five prompt‑optimisation algorithms and a single‑reflection baseline, especially when rollout costs are high.

By Yuval David, Fabiana Fournier, Lior Limonad, Hadar Mulian
arXiv Computation and Language
Aug 28

JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols

JudgeStealer is a query‑efficient framework that extracts the judging capabilities of large language models across pointwise scoring, pairwise comparison, and listwise ranking protocols. It leverages cross‑protocol agreement to convert pointwise scores into higher‑order supervision, dynamically selects informative inputs, and applies score smoothing and multi‑protocol review to preserve ordinal structure and avoid catastrophic forgetting. Experiments show it outperforms existing baselines, achieving up to 73.3% accuracy on pointwise, 87.0% on pairwise, and 71.6% on listwise evaluation, while remaining robust against common extraction defenses.

By Chen Chen, Yaolin Chen, Xuehan Sun, Juan Lin, Xueluan Gong, Yuhang Zheng, Qian Wang, Kwok-Yan Lam
arXiv AI
Jun 2

BADGER: Bridging Agentic and Deterministic Evaluation for Generative Enterprise Reasoning

arXiv:2606. 02109v1 Announce Type: new Abstract: Enterprise AI systems that translate natural language into SQL queries and orchestrate multi-step agentic reasoning pipelines require evaluation approaches fundamentally different from academic benchmarks.

By Shannon Serrao, Soumitra Chatterjee, Dorina Strori, Abhishek Sharma, Nathan Miller
arXiv AI
Sep 2

Autoresearch for Marketplace Catalogs: From Legacy Forms to AI-Native Matching

The paper describes a new approach for two‑sided service marketplaces that replaces fixed request forms with AI‑native probabilistic matching using large language models. It introduces an autoresearch loop that generates a provider‑side preference taxonomy for each occupation, iteratively refining candidate tag sets through a six‑rubric LLM judge and a seven‑critic panel. The system also maps legacy form questions back to the new taxonomy, enabling coverage assessment and human quality assurance.

By Kartik Ravisankar, Hojat Abdolanezhad, Daniel Capo, Sang Su Lee, Shishir Dash, Vijay Anand Raghavan
arXiv AI
Aug 13

RecSys Factory: Bounding LLM Agent Autonomy to Decision Points in the Industrial Recommender Lifecycle

arXiv:2608. 11241v1 Announce Type: new Abstract: Deploying LLM agents into industrial recommender operations exposes a three-way tension we frame as the autonomy-determinism-efficiency trilemma: general autonomy (interpreting operator intent, generating glue code zero-shot), industrial determinism (schema-conforming feature extraction, non-crashing A/B, zero compliance-path hallucination), and end-to-end efficiency.

By Dongyang Ao, Kaixiang Fang, Shijie Xu
arXiv AI
Jul 21

Fantastic Adaptive Taxonomies and How to Use Them

arXiv:2607. 16387v1 Announce Type: cross Abstract: An agent system's execution traces record how it fails, and procedures that improve such a system without changing model weights (trajectory selection, prompt and workflow optimization, runtime monitoring) read these traces for feedback.

By Mert Cemri, Andrei Cojocaru, Melissa Pan, Shu Liu, Shubham Agarwal, Alexander Krentsel, Jay Tang, Kannan Ramchandran, Joseph E. Gonzalez, Matei Zaharia, Alex Dimakis, Ion Stoica