arXiv Machine Learning

Repeated-Token Counting Reveals a Dissociation Between Representations and Outputs

arXiv:2605. 09239v2 Announce Type: replace-cross Abstract: Large language models fail at counting how many times a word repeats in a list, even though they perform well on far harder reasoning tasks.

arXiv Machine Learning
Jul 23

Reading Calibrated Uncertainty from Language Model Trajectories

arXiv:2605. 22864v2 Announce Type: replace Abstract: The maximum softmax probability (MSP) represents a default approach when evaluating uncertainty quantification for language model generation with structured output.

By Aliai Eusebi, Alexander Herzog, Xiaoyu Liang, Marie Vasek, Enrico Mariconti, Lorenzo Cavallaro