arXiv:2606. 30128v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting improves LLM reasoning, but the source is contested: do the intermediate steps help because they carry useful semantic content, or because conditioning on more tokens buys extra computation before the model commits to an answer?
By Wenlong Wang, Fergal Reid
arXiv:2606. 29490v1 Announce Type: cross Abstract: Confidence is an estimate of the probability that a chosen answer is correct.
By Dharshan Kumaran
arXiv:2606. 25013v1 Announce Type: new Abstract: Today's reasoning models use thinking tokens to attain stronger performance on benchmarks than their instruction-tuned counterparts.
By Narutatsu Ri, Abhishek Panigrahi, Sanjeev Arora
arXiv:2606. 26502v1 Announce Type: new Abstract: Large reasoning models (LRMs) take longer on harder problems, just as humans do.
By Han-yu Wang
arXiv:2607. 08059v1 Announce Type: cross Abstract: Uncertainty quantification for visual language models (VLMs) conventionally targets the answer token distribution.
By Mayank Singal
arXiv:2603. 29025v3 Announce Type: replace-cross Abstract: Large language models fail when a salient surface cue conflicts with an unstated feasibility constraint.
By Yubo Li, Lu Zhang, Tianchong Jiang, Ramayya Krishnan, Rema Padman