arXiv AI By Aditya Kumaran, Rahul Singhal, Karime Maamari, Amine Mhedhbi, Pradyumna Tambwekar

HARDEN: Constrained Evolutionary Search for Harder, Answer-Preserving Evaluation Cases

Read the original on arXiv AI →

HARDEN is a constrained evolutionary search method that transforms existing evaluation cases into more challenging variants while preserving their expected outputs. It operates along domain‑specific complexity axes and enforces feasibility constraints such as task semantics, realism, and execution validity. Experiments on FinQA, PubMedQA, and ContractNLI with Qwen3.5 models show that HARDEN can reduce task‑model accuracy by an average of 22.7% and up to 49.9% compared to single‑pass baselines.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 2

BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution

arXiv:2606. 01286v1 Announce Type: cross Abstract: The rapid progress of frontier large language models has led to widespread benchmark saturation, limiting the ability of existing datasets to differentiate model capabilities or provide useful training signal.

By Yangzhen Wu, Aaron J. Li, Wenjie Ma, Li Cao, Ziheng Zhou, Mert Cemri, Shu Liu, Yuran Xiu, Chenxiao Yan, Haikun Zhao, Bin Yu, Ion Stoica, Dawn Song
arXiv AI
Sep 17

Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts

arXiv:2609. 18366v1 Announce Type: new Abstract: Reliable agent evaluation is complicated by automatic harness optimization, which repeatedly uses a released benchmark $B_{\mathrm{rel}}$ to guide a Proposer that edits prompts, memory, retrieval, tools, and control code around a fixed target agent.

By Guojun Zhu, Xunheng Huang, Peng Yin, Jiahui Xie, Sanguo Zhang, Doudou Zhou