arXiv AI

FaultLens: Learning Compact Behavioral Test Suites for Generated Operational Programs

FaultLens is a technique for generating compact behavioral test suites for programs produced by automated generators. It learns probe orderings from earlier program generations, combining a fault‑driven greedy component with a mutation‑independent diversity component to cover a wide range of probe families, cases, templates, and temporal bins. In experiments on twenty generated operational policies across four environments, a 32‑probe hybrid suite learned from early generations covered 99.0% of dynamically killable faults in later generations while using only 1.2–2.0% of the exhaustive test domain.

Hugging Face Trending Papers
Aug 27

FaultLens: Learning Compact Behavioral Test Suites for Generated Operational Programs

FaultLens is a method for creating compact behavioral test suites for generated operational programs, balancing thoroughness with cost. It executes a rich probe domain once, stores fault‑probe kill relations, and learns probe orderings from earlier program generations using a fault‑driven greedy component and a mutation‑independent diversity component. In evaluations across multiple environments and program generations, a 32‑probe hybrid suite achieved 99.0% coverage of dynamically killable faults while using only 1.2‑2.0% of the exhaustive domain, and improved macro coverage when a fault family was withheld from training.

arXiv Machine Learning
Jul 10

TTHE: Test-Time Harness Evolution

arXiv:2607. 08124v1 Announce Type: cross Abstract: The behavior of an LLM agent is determined not only by the underlying model, but also by its harness: the executable program that constructs context, invokes tools, verifies intermediate results, and recovers from failures.

By Jun Nie, Yonggang Zhang, Jun Song, Qianshu Cai, Dahai Yu, Yike Guo, Xinmei Tian, Bo Han
Hugging Face Trending Papers
Jun 10

Layer-Isolated Evaluation: Gating the Deterministic Scaffold of a Production LLM Agent with a No-LLM, Regression-Locked Test Harness

End-to-end task-success is the dominant way to evaluate LLM agents, but one aggregate number tells you that an agent regressed, not where. We present layer-isolated evaluation: a deployed ordering agent is decomposed into a fixed taxonomy of layers (ontology, intent, routing, decomposition, escalation, safety, memory, and cross-cutting envelope/defense), each exercised by its own assertion slice in a deterministic, no-LLM "pure" mode.

arXiv AI
Jul 23

Beyond Fail-to-Pass: Iterative Hardening of Co-Generated Bug Reproduction Tests and Fixes

arXiv:2607. 19843v1 Announce Type: cross Abstract: Large language models (LLMs) have made automated program repair (APR) increasingly practical for real-world bugs, but repairing directly from bug reports remains underconstrained.

By Yuhao Tan, Zhibang Yang, Fangkai Yang, Yuan Yao, Yu Kang, Lu Wang, Pu Zhao, Xin Zhang, Xiaoxing Ma, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang
arXiv AI
Sep 18

Chronicle: Cut-Point Replay for Regression Testing of LLM Agents

Chronicle introduces a method called cut‑point replay to make regression testing of large language model (LLM) agents reproducible. It records an agent’s run at non‑deterministic boundaries as immutable envelopes and then replays selected boundaries while executing the rest live, enabling continuous‑integration tests that detect faulty code changes. Benchmarks show minimal overhead, perfect bit‑stability, and effective detection of unsafe actions in a mutation study.

By Tisha Chawla, Susheem Koul
arXiv AI
Aug 5

Don't Regenerate, Debug: A Domain-Specific Agent for Repairing Near-Miss Hardware Operators

arXiv:2608. 02712v1 Announce Type: cross Abstract: Kernel generation for hardware accelerators such as GPUs and NPUs has become a proving ground for large language models (LLMs), and state-of-the-art systems raise correctness through pipelines that couple LLMs with agentic reinforcement learning and evolutionary search.

By Yansong Sun, Shenxiu Wu, Siyuan Chen, Runlin Hou, Junhao Qiu, Junming Cao, Shudi Shao, Zhichao Lu, Qingfu Zhang