arXiv AI By Yubo Li, Ramayya Krishnan, Rema Padman

LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning

Read the original on arXiv AI →

arXiv:2608. 12321v1 Announce Type: cross Abstract: When a salient surface cue competes with an implicit feasibility constraint, LLMs often fail -- but aggregate accuracy conflates genuine constraint inference with conservative defaulting.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 22

Do Language Models Know Their Own Constraints?

The study investigates whether language models can explicitly report constraints they have learned through post‑training fine‑tuning. Using constrained recipe generation with five banned ingredients, the authors compare supervised fine‑tuning (SFT) and Group Relative Policy Optimization (GRPO) against an untrained baseline on a Constraint Awareness Benchmark. Both fine‑tuning methods increase behavioral compliance from 4% to about 90% but reduce explicit constraint reporting and erode retained third‑person knowledge, with GRPO showing more destructive effects. The results suggest that reward‑based signals may suppress constraints context‑independently, and that models fail to enumerate constraints on request even when they can avoid them internally.

By Arin Agarwal