arXiv Machine Learning By Lakshmana Sri Harsha Nemani, P. K. Srijith, Tomasz Ku\'smierczyk

Toward Efficient Uncertainty in LLMs through Evidential Knowledge Distillation

Read the original on arXiv Machine Learning →

arXiv:2507. 18366v2 Announce Type: replace Abstract: Accurate uncertainty quantification remains a key challenge for standard LLMs, prompting the adoption of Bayesian and ensemble-based methods.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 14

Reference-Based Distillation Detection in LLMs

arXiv:2607. 09692v1 Announce Type: new Abstract: Model distillation -- training on outputs from stronger third-party models -- is widely used to boost performance, but raises concerns about unfair advantages and policy violations.

By Rajat Rawat, Sizhe Chen, Akshay Anand, Michael Duan, Bob Rotsted, Sewon Min
arXiv Machine Learning
Aug 19

Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

The paper investigates why self‑distillation can sometimes worsen the reasoning abilities of large language models (LLMs). It finds that the process suppresses the model’s epistemic verbalization—its expression of uncertainty—leading to shorter but less accurate responses in mathematical reasoning tasks. Experiments on several LLMs show performance drops of up to 40%, especially on out‑of‑distribution problems where uncertainty expression is beneficial.

By Jeonghye Kim, Xufang Luo, Minbeom Kim, Sangmook Lee, Dohyung Kim, Jiwon Jeon, Dongsheng Li, Yuqing Yang
arXiv Machine Learning
1d ago

Distillation of Tabular Foundation Models into Efficient Predictors

The paper presents a method for distilling tabular foundation models (TFMs) into lightweight, dataset‑specific students. By using the full labeled training set as teacher context and training students on both observed and synthetic queries, the authors achieve significant performance gains over traditional supervised models on TabArena and TALENT benchmarks. The distilled students also provide substantial inference speedups, reducing the cost of repeated inference.

By Minho Jeong, Dooho Lee, Jinmo Lee, Jaemin Yoo