arXiv:2610.01428v1 Announce Type: cross
Abstract: Generalization in large language models (LLMs) is the ability to produce consistent and semantically stable outputs when the same input is expressed...
By Nagham Omar, Mahmoud Jabarin, Maya Rozenshtein, Rom Himelstein, Avi Mendelson, Amit LeVi
arXiv:2606. 06635v1 Announce Type: cross Abstract: Failures in language model reasoning emerge through distinct processes that leave identifiable signatures in the reasoning trace.
By Tanvi Thoria, Kiana Jafari, Marc R. Schlichting, Mykel J. Kochenderfer
arXiv:2606. 30653v1 Announce Type: cross Abstract: Large language models are increasingly deployed in agentic pipelines that depend on the model evaluating its own outputs without external verification.
By Marina Mancoridis, Zo\"e Hitzig
arXiv:2606. 18327v1 Announce Type: cross Abstract: Language models (LMs) that faithfully describe their own behavior can more easily be audited, understood, and trusted by users.
By Itamar Pres, Laura Ruis, Melat Ghebreselassie, Belinda Z. Li, Jacob Andreas
The paper investigates how different training strategies affect the prompt sensitivity of large language models. It reproduces and compares methods such as refined data construction and robustness objectives, finding that while robustness fine‑tuning improves over standard fine‑tuning and in‑context learning, the prompt gap remains large (40–57%). Notably, newer techniques like CoIN and PPCL often underperform a simple data‑construction approach that uses one template per batch, and diagnostics suggest that mixed‑template batches force the optimizer to reconcile conflicting updates rather than learn a prompt‑agnostic representation.
By Frederic Sadrieh, Michal \v{S}tef\'anik
arXiv:2507. 02778v3 Announce Type: replace-cross Abstract: Although large language models (LLMs) have transformed AI, they still make errors and follow unproductive reasoning paths.
By Ken Tsui