arXiv:2607. 16388v1 Announce Type: cross Abstract: Large-scale AI datacenter platforms comprise thousands of heterogeneous hardware components whose validation requires comprehensive fault injection test plans.
By Mohammed-Khalil Ghali, Saurabh Kulkarni, Prathamesh Kulkarni, Rohan Kulkarni, Sangwon Yoon, Daehan Won
arXiv:2606. 01008v1 Announce Type: cross Abstract: We present a benchmark for evaluating AI models and agents on real-world formal software verification tasks.
By Quinn Dougherty, Max von Hippel, Hazel Shackleton, Mike Dodds
arXiv:2609.39568v1 Announce Type: cross
Abstract: Large language models (LLMs) may generate unreliable code on corner cases missed by testing, while formal verification can provide machine-checkable...
By Jiaru Qian, Yihong Dong, Yongmin Li, Hao Zhu, Bin Gu, Ge Li
The paper introduces ReviveBench, a benchmark designed to evaluate coding agents’ ability to revive non‑running software and reconstruct industrial engines from open specifications. It comprises two families of tasks—revival (ten tasks addressing dependency issues, missing modules, legacy builds, and GPU models) and reconstruction (thirteen tasks covering numerical, geometric, hardware, and transactional systems). The benchmark uses hidden verifiers calibrated against native environments, engineering tools, or reference implementations, and the authors report that the strongest evaluated model passes all revival tasks and most reconstruction tasks, while also uncovering verifier defects that highlight measurement error in executable verification.
By Tianyu Liu, Dingyuan Dai, Yufan Du, Zhen Yang
arXiv:2609.22664v1 Announce Type: cross
Abstract: Research on large language model agents for penetration testing is evaluated almost entirely by capability: whether the agent captures a flag or repr...
By Joas Antonio dos Santos Barbosa
arXiv:2608. 19674v1 Announce Type: cross Abstract: Computing has been an astonishing success - but the accumulated technical debt exposes us all to huge costs in business and societal risk.
By Peter Sewell, Jean Pichon-Pharabod