arXiv AI By Xue Qin, Simin Luan, Cong Yang, Zhijun Li

Regression Test Selection for Updated Capability Modules in Compositional ML Systems via Atomic-Quality Probes

Read the original on arXiv AI →

arXiv:2604. 26689v4 Announce Type: replace-cross Abstract: Compositional machine-learning (ML) systems assemble runtime behavior from libraries of independently re-trained capability modules.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jun 10

Layer-Isolated Evaluation: Gating the Deterministic Scaffold of a Production LLM Agent with a No-LLM, Regression-Locked Test Harness

End-to-end task-success is the dominant way to evaluate LLM agents, but one aggregate number tells you that an agent regressed, not where. We present layer-isolated evaluation: a deployed ordering agent is decomposed into a fixed taxonomy of layers (ontology, intent, routing, decomposition, escalation, safety, memory, and cross-cutting envelope/defense), each exercised by its own assertion slice in a deterministic, no-LLM "pure" mode.