arXiv:2609.37841v1 Announce Type: new
Abstract: Masked generative models offer parallel token prediction, but accurate parallel sampling must account for dependencies among tokens. When dependencies...
By Ryotaro Kawata, Satoshi Hayakawa, Taiji Suzuki
arXiv:2609.13714v1 Announce Type: new
Abstract: An updated model can improve an aggregate metric while degrading a slice that matters to a downstream user. We study checkpoint selection subject to no...
By Shengwei Zhang, Tao Wu, Fei Qian
The paper introduces a risk‑controlled framework for using large language models (LLMs) as judges in tasks without reference answers. By calibrating uncertainty thresholds on a held‑out set, the method ensures that the false discovery rate of accepted verdicts stays below a user‑specified level α with high probability, using finite‑sample Clopper–Pearson intervals. When the parametric judge lacks confidence, the instance is routed to a retrieval‑augmented mode with a second calibrated threshold, preserving the error guarantee while achieving higher coverage than single‑mode baselines.
arXiv:2606. 29054v1 Announce Type: new Abstract: Large language models (LLMs) deployed for structured generation (NER, JSON extraction, QA, and classification) lack formal reliability guarantees, and standard heuristic abstention policies miss user-specified risk targets by 7.
By Varun Kotte
The paper critiques current memorization audits for generative models, arguing that lacking a proper null distribution leads to misleading conclusions. It introduces two exact null tests—one permutation test for whole models and a calibrated test for single images—showing that many previously flagged memorizations disappear under these stricter controls. The authors also propose a scale‑restricted statistic based on the Intersection Euler Characteristic Profile to better detect distinct copied images.
By Sushovan Majhi, Pramita Bagchi
The paper evaluates hard‑gate candidacy for validators in a deployed generative‑agent system by measuring how well each validator’s firing separates successful from failed builds. Across 13 validators and thousands of builds, only a few checks show statistically significant separation, while many fail to distinguish or never fire. The study highlights that skipped checks are recorded as passes, limiting detectable failure rates and underscoring the need for clearer evaluation records.
By Xin Xu
arXiv:2609.09245v1 Announce Type: new
Abstract: Repeated-sampling evaluations increasingly extrapolate pass@k far beyond the number n of samples collected per problem. We show that, in the pooled/ran...
By Pranav Singh, Prashant Singh
Repeated-sampling evaluations increasingly extrapolate pass@k far beyond the number n of samples collected per problem. We show that, in the pooled/random-task conditional-Binomial model, fixed-n succ...
The paper investigates how AI systems that perform best‑of‑n search require different validation strategies as the search width changes. It shows that auditing only small search widths leaves a gap in reliability estimates for larger widths, and proposes retaining candidate ranks and truth labels to estimate reliability across all widths up to N. The authors derive theoretical bounds on the minimax mean‑squared error, design procedures that achieve these bounds, and demonstrate that a shared audit can significantly reduce maximum error across many widths in practical CodeRM pools.
By Ricardo Fitas
arXiv:2607. 12649v1 Announce Type: new Abstract: Recent work on extractable memorization in LLMs suffers from two contrasting validity problems.
By A. Feder Cooper, Marika Swanberg, Jamie Hayes, Lea Duesterwald, Christopher De Sa, Daniel E. Ho, Mark A. Lemley, Percy Liang
arXiv:2608. 16852v1 Announce Type: new Abstract: Regulatory compliance monitoring in deployed language models is increasingly implemented as a legal and audit control, checking model outputs against written rules spanning data protection, healthcare, financial regulation, and platform policy.
By Saisab Sadhu, Aadit Sengupta, Vinay Kumar Sankarapu, Pratinav Seth
arXiv:2607. 19442v1 Announce Type: cross Abstract: Machine unlearning is commonly evaluated by matching a retrained oracle on trained probes.
By Sen Yang, Yuen-Hei Yeung