arXiv Machine Learning By Tezan Sahu, Aritra Das, Pankaj Mittal, Sudipta Das

Who Belongs in the Eval Set? A Capability-Taxonomy-Driven Pipeline for Curating Regression Eval Sets in Agent-Extensibility Platforms

Read the original on arXiv Machine Learning →

arXiv:2608. 01004v1 Announce Type: new Abstract: Platform teams hosting agent-extensibility surfaces face a regression-economics paradox: every onboarding customer ships an evaluation set tuned to their domain, but the platform's regression set must live under a hard query-count ceiling bounded by release cadence.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 18

Semantic Feature Analysis: Improving Agents Without Searching Over Rollouts

Semantic Feature Analysis (SFA) is a method that refines agent specifications without performing any rollout-based search. It analyzes existing execution traces, clusters workflow node outputs, extracts semantic feature classes via an extended subject‑verb‑object schema, ranks these features with a decision tree, and injects the most impactful features back into the system prompt. Evaluations on four benchmarks show that SFA consistently outperforms five prompt‑optimisation algorithms and a single‑reflection baseline, especially when rollout costs are high.

By Yuval David, Fabiana Fournier, Lior Limonad, Hadar Mulian