Hugging Face Trending Papers

FaultLens: Learning Compact Behavioral Test Suites for Generated Operational Programs

Read the original on Hugging Face Trending Papers →

FaultLens is a method for creating compact behavioral test suites for generated operational programs, balancing thoroughness with cost. It executes a rich probe domain once, stores fault‑probe kill relations, and learns probe orderings from earlier program generations using a fault‑driven greedy component and a mutation‑independent diversity component. In evaluations across multiple environments and program generations, a 32‑probe hybrid suite achieved 99.0% coverage of dynamically killable faults while using only 1.2‑2.0% of the exhaustive domain, and improved macro coverage when a fault family was withheld from training.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Aug 28

FaultLens: Learning Compact Behavioral Test Suites for Generated Operational Programs

FaultLens is a technique for generating compact behavioral test suites for programs produced by automated generators. It learns probe orderings from earlier program generations, combining a fault‑driven greedy component with a mutation‑independent diversity component to cover a wide range of probe families, cases, templates, and temporal bins. In experiments on twenty generated operational policies across four environments, a 32‑probe hybrid suite learned from early generations covered 99.0% of dynamically killable faults in later generations while using only 1.2–2.0% of the exhaustive test domain.

By Zeming Liu, Hang Lyu, Jingtao Zhang
arXiv Machine Learning
Jul 10

TTHE: Test-Time Harness Evolution

arXiv:2607. 08124v1 Announce Type: cross Abstract: The behavior of an LLM agent is determined not only by the underlying model, but also by its harness: the executable program that constructs context, invokes tools, verifies intermediate results, and recovers from failures.

By Jun Nie, Yonggang Zhang, Jun Song, Qianshu Cai, Dahai Yu, Yike Guo, Xinmei Tian, Bo Han
Hugging Face Trending Papers
Jun 10

Layer-Isolated Evaluation: Gating the Deterministic Scaffold of a Production LLM Agent with a No-LLM, Regression-Locked Test Harness

End-to-end task-success is the dominant way to evaluate LLM agents, but one aggregate number tells you that an agent regressed, not where. We present layer-isolated evaluation: a deployed ordering agent is decomposed into a fixed taxonomy of layers (ontology, intent, routing, decomposition, escalation, safety, memory, and cross-cutting envelope/defense), each exercised by its own assertion slice in a deterministic, no-LLM "pure" mode.