arXiv:2606. 20128v1 Announce Type: cross Abstract: Benchmarks for LLM-generated GPU kernels (KernelBench, TritonBench, GEAK) score correctness through fixed-shape, small-sample allclose-style checks.
By Dipankar Sarkar
The paper introduces mutation analysis as a metric for evaluating GPU‑kernel benchmark oracles, injecting over ten thousand faults into verified CUDA implementations of 188 KernelBench problems. It shows that the current official checkers miss 16.9% of faults, with precision faults being especially problematic, and demonstrates that optimized test suites can achieve 98% detection with only two inputs per problem. The study also reveals flaws in existing patches and a fuzzing recipe that incorrectly rejects correct kernels 107 times.
By Mingzhe Du, Anh Tuan Luu, Dong Huang, See-Kiong Ng
arXiv:2607. 16228v1 Announce Type: new Abstract: Most tensor-kernel correctness tests go through a fixed-shape all close-style check with hand-picked absolute and relative tolerances.
By Dipankar Sarkar
arXiv:2608. 12700v1 Announce Type: new Abstract: Systems that generate GPU kernels with language models report high correctness rates.
By Rishi Shah, Rishav Shrestha
arXiv:2606. 20502v1 Announce Type: cross Abstract: Whether LLMs scoring well on vulnerability benchmarks genuinely reason about security or merely pattern-match on contaminated data remains unresolved.
By Arastoo Zibaeirad, Marco Vieira
arXiv:2608. 08722v1 Announce Type: cross Abstract: Benchmarks for systems that are optimized against the evaluation signal measure something different from what they claim.
By V\'ictor Gallego
The paper investigates whether the hidden activations of large language models (LLMs) contain signals about the vulnerability of C/C++ code when the code is provided as context. By extracting prefill token activations from four LLMs and training small MLP probes, the authors achieve an average F1 score of 41.7% across four benchmarks, with the best probe matching state‑of‑the‑art fine‑tuned classifiers on the Devign dataset. The results suggest that a coding LLM’s internal representation can inform vulnerability detection, opening the door to lightweight, model‑native screening methods.
By Alizishaan Khatri
FaultLens is a technique for generating compact behavioral test suites for programs produced by automated generators. It learns probe orderings from earlier program generations, combining a fault‑driven greedy component with a mutation‑independent diversity component to cover a wide range of probe families, cases, templates, and temporal bins. In experiments on twenty generated operational policies across four environments, a 32‑probe hybrid suite learned from early generations covered 99.0% of dynamically killable faults in later generations while using only 1.2–2.0% of the exhaustive test domain.
By Zeming Liu, Hang Lyu, Jingtao Zhang
arXiv:2609. 19611v1 Announce Type: cross Abstract: Tensor programs, as used in deep learning models, are a prime target for optimization, as small performance improvements can have a large impact across training or inference workloads.
By Paul Biberstein, Joseph Devietti, Mayur Naik
arXiv:2609.23889v1 Announce Type: cross
Abstract: Automated kernel vulnerability reproduction is essential for bug triage, patch validation, and regression testing, but
still lacks an effective and...
By Xingyu Li, Juefei Pu, Haonan Li, Arrdya Srivastav, Kareem Shehada, Srikanth V. Krishnamurthy, Zhiyun Qian
arXiv:2605.01699v4 Announce Type: replace
Abstract: Recent attacks show that behavioural unlearning of large language models leaves internal traces recoverable by adversarial probes. We characterise...
By Anamika Paul Rupa, Anietie Andy
FaultLens is a method for creating compact behavioral test suites for generated operational programs, balancing thoroughness with cost. It executes a rich probe domain once, stores fault‑probe kill relations, and learns probe orderings from earlier program generations using a fault‑driven greedy component and a mutation‑independent diversity component. In evaluations across multiple environments and program generations, a 32‑probe hybrid suite achieved 99.0% coverage of dynamically killable faults while using only 1.2‑2.0% of the exhaustive domain, and improved macro coverage when a fault family was withheld from training.