Validation and Simulation Catch Different Errors: Four Levels of Evaluation for LLM-Generated Circuits
Read the original on arXiv AI →The paper introduces four distinct evaluation levels—schema validity, topological validity, backend executability, and component‑set agreement—to assess large language model‑generated electrical circuits. Using a 150‑circuit trilingual benchmark and a typed circuit interchange pipeline, the authors show that each level captures errors missed by the others, with significant discrepancies observed between validator rejections and ngspice execution outcomes. A repair study further demonstrates that targeted model adjustments can markedly improve topological validity while having mixed effects on executability and component agreement.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.