arXiv AI By Bryce Little

Length Penalties Make Chain-of-Thought Less Monitorable

Read the original on arXiv AI →

arXiv:2607. 09786v1 Announce Type: new Abstract: Length-penalized reinforcement learning can shorten chain-of-thought reasoning while hiding an influence that drives the model's answer.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.