Hugging Face Trending Papers

FaultLens: Learning Compact Behavioral Test Suites for Generated Operational Programs

FaultLens is a method for creating compact behavioral test suites for generated operational programs, balancing thoroughness with cost. It executes a rich probe domain once, stores fault‑probe kill relations, and learns probe orderings from earlier program generations using a fault‑driven greedy component and a mutation‑independent diversity component. In evaluations across multiple environments and program generations, a 32‑probe hybrid suite achieved 99.0% coverage of dynamically killable faults while using only 1.2‑2.0% of the exhaustive domain, and improved macro coverage when a fault family was withheld from training.

arXiv AI
Aug 28

FaultLens: Learning Compact Behavioral Test Suites for Generated Operational Programs

FaultLens is a technique for generating compact behavioral test suites for programs produced by automated generators. It learns probe orderings from earlier program generations, combining a fault‑driven greedy component with a mutation‑independent diversity component to cover a wide range of probe families, cases, templates, and temporal bins. In experiments on twenty generated operational policies across four environments, a 32‑probe hybrid suite learned from early generations covered 99.0% of dynamically killable faults in later generations while using only 1.2–2.0% of the exhaustive test domain.

By Zeming Liu, Hang Lyu, Jingtao Zhang
arXiv Machine Learning
Jul 10

TTHE: Test-Time Harness Evolution

arXiv:2607. 08124v1 Announce Type: cross Abstract: The behavior of an LLM agent is determined not only by the underlying model, but also by its harness: the executable program that constructs context, invokes tools, verifies intermediate results, and recovers from failures.

By Jun Nie, Yonggang Zhang, Jun Song, Qianshu Cai, Dahai Yu, Yike Guo, Xinmei Tian, Bo Han
Hugging Face Trending Papers
Jun 10

Layer-Isolated Evaluation: Gating the Deterministic Scaffold of a Production LLM Agent with a No-LLM, Regression-Locked Test Harness

End-to-end task-success is the dominant way to evaluate LLM agents, but one aggregate number tells you that an agent regressed, not where. We present layer-isolated evaluation: a deployed ordering agent is decomposed into a fixed taxonomy of layers (ontology, intent, routing, decomposition, escalation, safety, memory, and cross-cutting envelope/defense), each exercised by its own assertion slice in a deterministic, no-LLM "pure" mode.

arXiv AI
Jul 23

Beyond Fail-to-Pass: Iterative Hardening of Co-Generated Bug Reproduction Tests and Fixes

arXiv:2607. 19843v1 Announce Type: cross Abstract: Large language models (LLMs) have made automated program repair (APR) increasingly practical for real-world bugs, but repairing directly from bug reports remains underconstrained.

By Yuhao Tan, Zhibang Yang, Fangkai Yang, Yuan Yao, Yu Kang, Lu Wang, Pu Zhao, Xin Zhang, Xiaoxing Ma, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang