Sebastian Raschka By Sebastian Raschka, PhD

GPT-6 Astra, Looped Transformers, and Hidden Reasoning

Read the original on Sebastian Raschka →

The article titled "GPT-6 Astra, Looped Transformers, and Hidden Reasoning" examines recent developments in transformer architecture, focusing on recurrent depth, hidden chains of thought, and the concept of looping transformer blocks. It discusses how these innovations aim to enhance the reasoning capabilities of language models by allowing deeper, more iterative processing of information. The piece highlights current research trends that explore the potential of these techniques to improve model performance and interpretability.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Sebastian Raschka.

arXiv Machine Learning
5d ago

Looped Transformers as Optimizers

arXiv:2609.37379v1 Announce Type: new Abstract: Looped Transformers provide a parameter-efficient approach to depth scaling by repeatedly applying shared Transformer blocks. Recent reasoning models h...

By Yulong Huang, Chen Jiang, Zhanpeng Zhou, Hongtao Zhang, Tianyu Li, Tianyu He, Xiangyu Zhang, Bojun Cheng
arXiv Machine Learning
2d ago

Decoding Looped Transformers Better for (Almost) Free

The paper introduces LoopCD, a training‑free contrastive decoding framework that improves token selection in Loop‑Transformer models by comparing the final prediction with earlier recurrent passes. LoopCD operates either in logit space (LoopCD‑Logits) with a single extra output pass or in hidden‑state space (LoopCD‑Hidden) with no output overhead. Across multiple looped Transformer families, LoopCD yields significant performance gains—raising pass@1 scores on tasks such as AIME 2024 and HumanEval—while enabling a reduction in the number of recurrent loops and a corresponding decrease in inference FLOPs.

By Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang