The paper introduces mutation analysis as a metric for evaluating GPU‑kernel benchmark oracles, injecting over ten thousand faults into verified CUDA implementations of 188 KernelBench problems. It shows that the current official checkers miss 16.9% of faults, with precision faults being especially problematic, and demonstrates that optimized test suites can achieve 98% detection with only two inputs per problem. The study also reveals flaws in existing patches and a fuzzing recipe that incorrectly rejects correct kernels 107 times.
By Mingzhe Du, Anh Tuan Luu, Dong Huang, See-Kiong Ng
arXiv:2606. 16062v1 Announce Type: new Abstract: We measure the rate at which code RL environments accept incorrect solutions as correct.
By Shreshth Rajan
arXiv:2605.01699v4 Announce Type: replace
Abstract: Recent attacks show that behavioural unlearning of large language models leaves internal traces recoverable by adversarial probes. We characterise...
By Anamika Paul Rupa, Anietie Andy
arXiv:2606. 10229v1 Announce Type: cross Abstract: We study whether demonstration-curation metrics that detect defective training episodes also improve the downstream behavior-cloning policy that trains on the curated data.
By Aarav Bedi
arXiv:2606. 20502v1 Announce Type: cross Abstract: Whether LLMs scoring well on vulnerability benchmarks genuinely reason about security or merely pattern-match on contaminated data remains unresolved.
By Arastoo Zibaeirad, Marco Vieira
arXiv:2604. 17388v3 Announce Type: replace-cross Abstract: Time series anomaly detectors have grown steadily more complex, incorporating attention mechanisms, adversarial training, and stochastic latent variables.
By Kadir-Kaan \"Ozer, Ren\'e Ebeling, Markus Enzweiler
arXiv:2607. 19442v1 Announce Type: cross Abstract: Machine unlearning is commonly evaluated by matching a retrained oracle on trained probes.
By Sen Yang, Yuen-Hei Yeung
arXiv:2606. 27396v1 Announce Type: cross Abstract: Test-input generation for tensor kernels is folkloric.
By Dipankar Sarkar
arXiv:2607. 11022v1 Announce Type: new Abstract: The test suites used as RLVR rewards for code have natural false positives: per-task, persistent, asymmetric errors that accept the same wrong programs every time they appear, unlike the symmetric or resampled noise assumed by existing noise-robustness analyses.
By Chuyifei Zhang
arXiv:2607. 07146v1 Announce Type: new Abstract: The Internal Waves Service screens the Sentinel-1 Wave-mode archive for internal solitary waves, routing detections to experts whose adjudication time is the resource the effort exists to conserve.
By Joao Pinelo, Joao Goncalves, Arun Shukla, Adriana Santos-Ferreira
arXiv:2609.21801v1 Announce Type: new
Abstract: We study how far a simple statistical pipeline can go on univariate time series anomaly detection under a strict selection protocol. The method extract...
By Youssef Attia El Hili, Malik Tiomoko, Corinne Ancourt
arXiv:2608. 08008v1 Announce Type: new Abstract: Process reward models (PRMs) score intermediate reasoning steps and are widely used for search, ranking, and training, but optimization can exploit these learned proxies by increasing reward while turning correct reasoning into incorrect reasoning.
By Ibne Farabi Shihab, Fariya Afrin