arXiv Machine Learning By Hanti Lin

Cross-Entropy Risk Estimation for Language Models: Inconsistency Must Be Dense, and the Holdout Method Is No Exception

Read the original on arXiv Machine Learning →

arXiv:2608. 15798v1 Announce Type: new Abstract: Language models are compared by their held-out per-token cross-entropy risk---the quantity scaling laws are fitted to.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.