Microsoft Research

Verifying Rust cryptography in SymCrypt, from standards to code

Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves.

arXiv AI
Jul 7

RustMizan: A Compilable, Contamination-Aware Benchmarking Framework for Rust Vulnerabilities

arXiv:2607. 04729v1 Announce Type: cross Abstract: LLM agents are increasingly applied to vulnerability analysis, but existing benchmarks have not kept pace.

By Tarek Elsayed, Shiping Yang, Eunsong Koh, Sanika Goyal, Vincent Huang, Paul Ngo, Nathan Young, Mohammad Omidvar Tehrani, Alvyn Kang, Arnell Kang, Zeyu Chen, Ang\'elica Moreira, Xuan Feng, Angel X. Chang, Nick Sumner, Steven Y. Ko
arXiv AI
Sep 10

Scratchy: Visual-Scratchpad Multimodal Reasoning for Cryptographic Proof Generation in EasyCrypt

Scratchy introduces a visual-scratchpad method for generating cryptographic proofs in EasyCrypt by converting natural-language security descriptions into a typed proof-relation graph and then into a visual proof state that guides multimodal language models. The approach exposes implicit proof-theoretic dependencies that LLMs struggle with, enabling clearer coordination of probability, adversarial games, invariants, assumptions, and bounds. Scratchy-eval, a 114-task dataset from official EasyCrypt files, demonstrates that classical LLMs achieve significant gains when using these structured visual proof states.

By Yupeng Ren, Zhaoxuan Li, Rui Zhang
OpenAI Blog
Dec 18, 2025

Introducing GPT-5.2-Codex

GPT-5. 2-Codex is OpenAI’s most advanced coding model, offering long-horizon reasoning, large-scale code transformations, and enhanced cybersecurity capabilities.

arXiv Machine Learning
Jul 27

Certified in Theory, Broken in Practice: Assumption Gaps in Cryptographic Model Certification

arXiv:2607. 21839v1 Announce Type: cross Abstract: Privacy-preserving machine learning auditing protocols allow auditors to assess models for properties such as accuracy or fairness, without revealing their internals or training data.

By Carter Luck, Olive Franzese-McLaughlin, Elisaweta Masserova, Akira Takahashi, Antigoni Polychroniadou, Nicolas Papernot