arXiv:2607. 21372v1 Announce Type: cross Abstract: Score Entropy Discrete Diffusion (SEDD) parameterizes discrete reverse processes with unconstrained positive score ratios.
By Jingyuan Li, Xiaoyi Jiang, Yixuan Jiang, Wei Liu, Yi Zhu, Zuoqiang Shi, Pipi Hu
arXiv:2607. 24507v1 Announce Type: cross Abstract: Existing methods mainly adapt pretrained autoregressive (AR) language models to masked diffusion, whereas we directly adapt them to uniform-noise diffusion, where every token remains editable during sampling.
By Xiaoyi Jiang, Jingyuan Li, Yixuan Jiang, Wei Liu, Yi Zhu, Zuoqiang Shi, Pipi Hu
arXiv:2609.39934v1 Announce Type: cross
Abstract: Checkpoint selection in domain generalization often relies on source-validation accuracy, yet the selected checkpoint need not provide reliable proba...
By Jinshi Liu, Jiahao Li, Pan Liu, Yanfeng Li, Rui Qian, Zhao Tong, Yue Sun, Tao Tan
arXiv:2605. 23434v2 Announce Type: replace Abstract: Approximate inference over inducing variables is the central computational bottleneck of Deep Gaussian Processes (DGPs).
By Jian Xu, Delu Zeng, John Paisley, Qibin Zhao
arXiv:2608. 01023v1 Announce Type: new Abstract: We present Caliber, an output-perturbation defense against model extraction that formulates noise selection as a calibration problem: how much the defense degrades the supervision signal used to train a surrogate, and the provable per-input query cost of recovering the clean logits.
By Chi Wang, Hanwen Wang, Yu Xia, Zihan Wang, Guangdong Bai
arXiv:2608. 14647v1 Announce Type: cross Abstract: Dirty-history rollouts make multi-turn on-policy self-distillation (OPSD) brittle: once a student emits an erroneous intermediate reply, later turns are conditioned on that reply, and uniform distillation can spend loss on tokens that carry little corrective signal.
By Chenyang Jiang, Changhan Huang
arXiv:2607. 05381v1 Announce Type: cross Abstract: What does a discrete diffusion model learn: a denoiser, a score ratio, or a bridge plug-in predictor?
By Rodrigo Casado Noguerales, Bernhard Sch\"olkopf, Thomas Hofmann, Aran Raoufi
arXiv:2608.23849v1 Announce Type: new
Abstract: Negative sampling determines whether a knowledge graph embedding (KGE) model learns from informative counterexamples or wastes updates on implausible c...
By Ibne Farabi Shihab, Naoshin Anzum Hridi, Joyanta Jyoti Mondal
arXiv:2607. 13124v1 Announce Type: cross Abstract: Structured pruning is a hardware-friendly way to compress LLMs, but it is mostly validated on multiple-choice recognition tasks, while the same compressed checkpoints can collapse on the free-form generation that deployment actually requires.
By Qingyu Zhang, Qianhao Yuan, Hongyu Lin, Yaojie Lu, Xianpei Han, Le Sun, Xiang Li, Ming Xu, Jiarui Li, Xiuyin Zhao
The paper introduces committed reveal sampling (CRS), a training‑free sampler for uniform discrete diffusion models that stores selected argmax tokens as persistent context for subsequent predictions. CRS keeps these tokens visible in later model inputs, which theoretically prevents Bayes error from increasing as noise decreases and encourages consistent sequence‑level choices. Empirical tests on Duo‑distilled data show that CRS without top‑p truncation achieves lower generative perplexity than fixed‑p baselines across various numbers of function evaluations, offering a more favorable perplexity–entropy trade‑off.
By Satoshi Hayakawa
arXiv:2607. 28864v1 Announce Type: cross Abstract: Tree-based diffusion models fit flexible conditional predictive distributions for tabular regression without a neural density estimator, but they inherit their design defaults---noising path, parameterization, training distribution, features, sampler---from the neural setting.
By Silas Koemen
Negative sampling determines whether a knowledge graph embedding (KGE) model learns from informative counterexamples or wastes updates on implausible corruptions. Uniform negatives are diverse but eas...