arXiv Machine Learning By Itai Morad, Nir Shlezinger, Yonina C. Eldar

SGD-Based Knowledge Distillation with Bayesian Teachers: Theory and Guidelines

Read the original on arXiv Machine Learning →

arXiv:2601. 01484v2 Announce Type: replace Abstract: Knowledge Distillation (KD) is a central paradigm for transferring knowledge from a large teacher network to a typically smaller student model, often by leveraging soft probabilistic outputs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 29

Rethinking Likelihood distributions: Student's t Likelihood Boosts Bayesian Neural Network Performance

arXiv:2607. 25376v1 Announce Type: cross Abstract: In Bayesian neural networks (BNNs), variational inference is a widely adopted framework for modeling uncertainty in a distributional way, with the evidence lower bound (ELBO) serving as the standard objective function.

By Pei-Hsuan Hsia, Lars H. Heyen, Arvid Weyrauch, Markus Goetz, Achim Streit, Sebastian Krumscheid, Charlotte Debus