arXiv AI

Mean-to-Score Discrete Diffusion: Posterior-Mean Denoisers for Score Entropy

arXiv:2607. 21372v1 Announce Type: cross Abstract: Score Entropy Discrete Diffusion (SEDD) parameterizes discrete reverse processes with unconstrained positive score ratios.

arXiv AI
Jul 28

UNIFUSION: Adapting Autoregressive Language Models into Discrete Diffusion under a Unified Reverse-Rate Objective

arXiv:2607. 24507v1 Announce Type: cross Abstract: Existing methods mainly adapt pretrained autoregressive (AR) language models to masked diffusion, whereas we directly adapt them to uniform-noise diffusion, where every token remains editable during sampling.

By Xiaoyi Jiang, Jingyuan Li, Yixuan Jiang, Wei Liu, Yi Zhu, Zuoqiang Shi, Pipi Hu
arXiv Machine Learning
Aug 4

Caliber: Cross-Architecture Extraction-Cost Control for Score-Returning APIs

arXiv:2608. 01023v1 Announce Type: new Abstract: We present Caliber, an output-perturbation defense against model extraction that formulates noise selection as a calibration problem: how much the defense degrades the supervision signal used to train a surrogate, and the provable per-input query cost of recovering the clean logits.

By Chi Wang, Hanwen Wang, Yu Xia, Zihan Wang, Guangdong Bai
arXiv Machine Learning
1d ago

Sharpen Before You Adapt: Data-Free Entry-State Sharpening for Test-Time Reinforcement Learning

The paper introduces entry-state sharpening, a data‑free pre‑training step that prepares a language model’s checkpoint in a sharper, lower‑entropy state before test‑time reinforcement learning (TTRL). By reducing policy entropy, the model can more efficiently use its limited adaptation budget, leading to higher endpoint conversion efficiency across tasks such as MATH, GPQA, and AMC. Experiments with different data‑free objectives (e.g., R‑Zero vs. SPIRAL) demonstrate that the choice of pre‑training objective strongly influences the checkpoint’s readiness for TTRL, and a label‑free self‑distillation intervention can further sharpen the entry state.

By Zhanming Zhang, Vinoth Selvendran
arXiv AI
Aug 18

SMOPD: Selective Token-Entropy Masking for Dirty-History Multi-Turn On-Policy Self-Distillation

arXiv:2608. 14647v1 Announce Type: cross Abstract: Dirty-history rollouts make multi-turn on-policy self-distillation (OPSD) brittle: once a student emits an erroneous intermediate reply, later turns are conditioned on that reply, and uniform distillation can spend loss on tokens that carry little corrective signal.

By Chenyang Jiang, Changhan Huang