arXiv AI By Man Liang, Xinzhao Cheng, Faizan Wajid

Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints

Read the original on arXiv AI →

The paper audits frozen decoder‑only large language models (LLMs) on geometric reasoning tasks using parametric CAD constraints. It probes hidden states for linear decodability, forced‑choice generation, activation‑level influence, and behavioral steerability, finding that pretraining improves decoding of local geometric relations but not sketch‑level DOF status. The study shows that decodable information is not always actionable: generation often fails to express it, and steering interventions do not reliably control outputs, revealing divergences among decodability, generation, activation influence, and steerability.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 5

Reasoning Models Don't Just Think Longer, They Move Differently

arXiv:2605. 15454v2 Announce Type: replace-cross Abstract: Reasoning-trained language models often spend more tokens on harder problems, but longer chains of thought do not show whether a model is merely computing for more steps or following a different internal trajectory.

By Anders Gj{\o}lbye, Lars Kai Hansen, Sanmi Koyejo
arXiv AI
Sep 11

Geometry Conditioning in an Embodied SLM: Training Controls and Robustness Diagnostics in a 0.8B Hybrid Model

The paper investigates how incorporating physical-state (geometric) inputs influences a 0.8B hybrid language model with 6.2M trainable parameters for manipulation tasks. Training with geometry-conditioned recurrent decay gates achieves a 28.9% success rate, slightly lower than the 36.7% success when geometry increments are shuffled, and higher than the 24.4% success without explicit geometry. Robustness tests show that a state-only relative-coordinate policy retains most success under frame relabeling, whereas visual policies degrade significantly after small object displacements, indicating no clear advantage from training-time geometric alignment under the tested conditions.

By Hao Li, Haofei Sun, Lin He
arXiv AI
Jul 15

Visual Access Boundaries in Vision-Language Model Reasoning

arXiv:2607. 12815v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended when VLMs generate longer reasoning traces.

By Hiroto Osaka, Shohei Taniguchi, Gouki Minegishi, Kai Yamashita, Masahiro Suzuki, Yutaka Matsuo
Hugging Face Trending Papers
Jul 14

Visual Access Boundaries in Vision-Language Model Reasoning

Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended when VLMs generate longer reasoning traces. We ask whether CoT requires continued access to image tokens, or whether it mainly operates over visual information already made available earlier in the forward pass.