The paper investigates mechanistic interpretability, focusing on how automated circuit discovery is evaluated. It shows that the commonly used faithfulness objective can favor circuits that reproduce a model’s behavior poorly, creating an objective-level recovery gap. Experiments on four human-reference tasks and InterpBench reveal that many discovery methods misrank candidate circuits, and that restoring excluded signals can correct most of these misrankings without altering the circuits’ behavior.
By Chuqin Geng, Li Zhang, Haolin Ye, Mark Zhang, Luke Zhang, Xujie Si
The paper introduces four distinct evaluation levels—schema validity, topological validity, backend executability, and component‑set agreement—to assess large language model‑generated electrical circuits. Using a 150‑circuit trilingual benchmark and a typed circuit interchange pipeline, the authors show that each level captures errors missed by the others, with significant discrepancies observed between validator rejections and ngspice execution outcomes. A repair study further demonstrates that targeted model adjustments can markedly improve topological validity while having mixed effects on executability and component agreement.
By Ali Hedayati Pirouzan
arXiv:2608. 13754v1 Announce Type: new Abstract: The EU AI Act requires providers of high-risk systems to file technical documentation describing how the system reaches its decisions.
By Ajay Pravin Mahale (Hochschule Trier)
Simulation success is not equivalent to structural correctness for LLM-generated circuits. We define and measure four evaluation levels -- schema validity, topological validity, backend executability,...
arXiv:2608. 07914v1 Announce Type: new Abstract: Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power.
By Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma
arXiv:2609.23549v1 Announce Type: new
Abstract: Screening oxygen-evolution catalysts on combinatorial libraries requires deciding which candidates receive the remaining measurements. The deciding act...
By Yong-Woon Kim, Jihyeok Lee, Sungtae Park, Sooseok Choi, Yung-Cheol Byun
The paper presents a hierarchical deep‑learning framework that integrates compositional, structural, and transport models to screen solid‑state electrolytes. Four modules—L‑G‑DCNN, DenseGNN, MatterSim, and DeePMD—coordinate to evaluate thermodynamics, multi‑property performance, and kinetic transport, outperforming existing methods. Applied to over 30 million candidates, the workflow identifies 97 high‑performance materials, mainly halides, and links Li⁺ jump‑network connectivity to ionic conductivity while highlighting limits for oxide electrolytes.
By Hongwei Du, Dingyang Lv, Baole Wei, Yongheng Li, Feng Yu, Ziheng Lu, Siqi Shi, Hong Wang
arXiv:2607. 16864v1 Announce Type: new Abstract: Supercharging of lithium-ion batteries (LiBs) requires robust health monitoring to ensure durability, safety, and user confidence, particularly for emerging vehicle-to-grid applications with bidirectional energy flows.
By Wendi Guo, S{\o}ren Byg Vilsen, Daniel Ioan Stroe, Yaqi Li, Yicun Huang, Ashima Verma, Daniel Brandell
arXiv:2607. 18921v1 Announce Type: cross Abstract: Circuit extraction identifies a small set of model components whose presence preserves a target behavior under ablation, and the resulting circuit is often read as the mechanism behind that behavior.
By Yang Sheng, Jie Fu
arXiv:2608. 16212v1 Announce Type: new Abstract: Laboratory battery tests provide the main empirical basis for battery performance and degradation studies, but their operating patterns do not directly represent field duty profiles.
By Chunyang Zhao, Chresten Tr{\ae}holt
arXiv:2608. 12652v1 Announce Type: cross Abstract: Benchmark contamination is diagnosed today with n-gram overlap, with likelihood-based membership inference, or with canary strings, and each needs something usually unavailable: the training corpus, a well-chosen test statistic, or foresight at dataset release.
By Florian Braun
arXiv:2607. 09762v1 Announce Type: new Abstract: Public battery aging datasets are a critical asset for advanced health management, but their practical use is often limited by inconsistent formats, unclear schemas, and metadata scattered across repositories and publications.
By Tianwen Zhu, Hao Wang, Yonggang Wen