arXiv AI

Structural Certification for Reliable Physical Design with Language Models

arXiv:2606. 30107v1 Announce Type: new Abstract: An unreliable language model can be made to produce reliable physical designs if the authority to assert is moved out of the model: the model proposes, and a deterministic engine alone certifies, returning certified, impossible, or unknown.

arXiv AI
Aug 24

Prediction certification cannot replace explanation certification: a competence envelope for trustworthy AI under compound stress

The paper argues that prediction‑based certifications—such as accuracy, calibration, and conformal coverage—are insufficient to guarantee trustworthy AI. It proves a separation theorem showing that a model can appear reliable under all prediction‑side certificates yet differ arbitrarily in explanation fidelity and deployment behaviour. The authors propose a competence envelope framework that combines both prediction and explanation certification to detect such hidden failures.

By Nataliya Shakhovska, Ivan Izonin, Stergios-Aristoteles Mitoulis
arXiv AI
Sep 30

Position: Let's Strengthen Verifiability If We Can't Enforce Reproducibility

The paper argues that empirical results in Machine Learning are often difficult to reproduce due to limited availability of code and supporting materials, which hampers research progress. It analyzes and quantifies these challenges and proposes concrete measures to enhance the verifiability of results, even if full reproducibility cannot be guaranteed. The authors provide their code and supporting resources on GitHub for reference.

By Samet Hicsonmez, Nermin Samet, Renaud Marlet
arXiv AI
Jul 7

Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models

arXiv:2607. 05199v1 Announce Type: new Abstract: Physics reasoning fails structurally in small language models: an error at any step propagates forward, corrupting every inference that follows.

By Raj Jaiswal, Dhruv Jain, Rishabh Dhawan, Sree Krishna Uppalapati, Shin'ichi Satoh, Tanuja Ganu, Rajiv Ratn Shah
arXiv AI
Sep 18

Trust, but Validate the Instrument: Auditing AI-Generated RTL Verification Plans on Authored Security-Regression Proxies

The paper introduces SecTB-RTL, an auditable framework for evaluating AI-generated RTL verification plans against 31 tasks and 124 hardware‑security regressions. In a confirmatory run, the AI model’s responses were rejected by the provider’s schema, and after a schema‑only repair, only nine of 1,857 accepted responses passed the production semantic validator, revealing a mismatch between generation and execution rules. The study demonstrates that schema acceptance does not guarantee execution validity and provides a benchmark, failure‑preserving contract, incident provenance, and governance controls to prevent misreporting of infrastructure behavior as model behavior.

By Hang Xiao, Chuhong Xu, Kainan Zhou, Gangzhen Qian, Lu Yi
Hugging Face Trending Papers
5d ago

Certification of Real Images through Calibrated Content Authentication

The paper examines the reliability of deepfake detectors, noting a decline in accuracy from 99.5% to 76% over four years and a drastic drop below 2% when adversarial perturbations are applied. It argues that content alone cannot determine provenance because generators can perfectly reproduce authentic images, so a new detection approach is proposed that certifies authenticity only if a generator cannot faithfully reconstruct the content. The authors demonstrate that their calibrated detector can limit false certifications to 1% while maintaining robustness against bounded‑perturbation attacks, though it struggles with arbitrary transformations.