Validation and Simulation Catch Different Errors: Four Levels of Evaluation for LLM-Generated Circuits
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The paper introduces four distinct evaluation levels—schema validity, topological validity, backend executability, and component‑set agreement—to assess large language model‑generated electrical circuits. Using a 150‑circuit trilingual benchmark and a typed circuit interchange pipeline, the authors show that each level captures errors missed by the others, with significant discrepancies observed between validator rejections and ngspice execution outcomes. A repair study further demonstrates that targeted model adjustments can markedly improve topological validity while having mixed effects on executability and component agreement.
The paper evaluates hard‑gate candidacy for validators in a deployed generative‑agent system by measuring how well each validator’s firing separates successful from failed builds. Across 13 validators and thousands of builds, only a few checks show statistically significant separation, while many fail to distinguish or never fire. The study highlights that skipped checks are recorded as passes, limiting detectable failure rates and underscoring the need for clearer evaluation records.
arXiv:2609.05881v1 Announce Type: cross Abstract: Developers increasingly run large language models locally by pulling quantized GGUF artifacts from public registries, yet nothing in the distribution...
arXiv:2607. 28871v1 Announce Type: cross Abstract: When a repair agent runs a test and sees it pass, the result is treated as evidence about the reported defect.
arXiv:2607. 06636v1 Announce Type: cross Abstract: Large language models frequently generate code that appears correct on typical inputs yet fails on edge cases, invalid inputs, and other specification-defined corner conditions.
The paper investigates the safety of Code World Models, where a language model generates executable world models that a planner uses. It shows that accepting a model based on sampled transitions only guarantees sample consistency, not full safety, because the probability of missing critical events decays as (1‑r)^N. Experiments on hybrid instruments reveal that omitted mode‑boundaries can severely limit planner performance, and that even sophisticated LLMs (GPT‑5.x) struggle to repair such omissions in higher‑dimensional settings.