arXiv:2608.30751v1 Announce Type: new
Abstract: Large language models (LLMs) trained only on text and code can sometimes generate programs that draw recognizable images. However, it is unclear whethe...
By Ashwin Nedungadi, Stefan Oehmcke, Stefan L\"udtke
arXiv:2605. 15454v2 Announce Type: replace-cross Abstract: Reasoning-trained language models often spend more tokens on harder problems, but longer chains of thought do not show whether a model is merely computing for more steps or following a different internal trajectory.
By Anders Gj{\o}lbye, Lars Kai Hansen, Sanmi Koyejo
Large language models (LLMs) trained only on text and code can sometimes generate programs that draw recognizable images. However, it is unclear whether this reflects an internal representation of 2D...
The paper investigates how incorporating physical-state (geometric) inputs influences a 0.8B hybrid language model with 6.2M trainable parameters for manipulation tasks. Training with geometry-conditioned recurrent decay gates achieves a 28.9% success rate, slightly lower than the 36.7% success when geometry increments are shuffled, and higher than the 24.4% success without explicit geometry. Robustness tests show that a state-only relative-coordinate policy retains most success under frame relabeling, whereas visual policies degrade significantly after small object displacements, indicating no clear advantage from training-time geometric alignment under the tested conditions.
By Hao Li, Haofei Sun, Lin He
arXiv:2607. 12815v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended when VLMs generate longer reasoning traces.
By Hiroto Osaka, Shohei Taniguchi, Gouki Minegishi, Kai Yamashita, Masahiro Suzuki, Yutaka Matsuo
Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended when VLMs generate longer reasoning traces. We ask whether CoT requires continued access to image tokens, or whether it mainly operates over visual information already made available earlier in the forward pass.