arXiv:2607. 18292v2 Announce Type: replace Abstract: As language models scale, answers start truer but degrade faster: scaling buys capability but erodes reliability.
By Kushal Chakrabarti
arXiv:2607. 18292v3 Announce Type: replace-cross Abstract: Bigger language models are less reliable.
By Kushal Chakrabarti
arXiv:2608. 20290v1 Announce Type: new Abstract: Whether a language model has improved itself is increasingly judged not by mean accuracy but by which individual problems it gains and loses.
By Cheng Xu, Nan Yan, Liming Chen, M-Tahar Kechadi
Whether a language model has improved itself is increasingly judged not by mean accuracy but by which individual problems it gains and loses. Tracking these transitions means differencing two noisy estimates, leaving them vulnerable to measurement artifacts.
arXiv:2610.00694v1 Announce Type: cross
Abstract: Compression reports summarize how far a compressed language model moved from the dense one, usually by a KL divergence; a deployment that relies on t...
By Beatriz Almeida Felicio
arXiv:2608. 13063v1 Announce Type: new Abstract: Prior work on LLM behavior under anomalous conditions asks whether a model notices anomalies.
By Sam Mao