arXiv AI

LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning

arXiv:2608. 12321v1 Announce Type: cross Abstract: When a salient surface cue competes with an implicit feasibility constraint, LLMs often fail -- but aggregate accuracy conflates genuine constraint inference with conservative defaulting.

arXiv Machine Learning
Sep 22

Do Language Models Know Their Own Constraints?

The study investigates whether language models can explicitly report constraints they have learned through post‑training fine‑tuning. Using constrained recipe generation with five banned ingredients, the authors compare supervised fine‑tuning (SFT) and Group Relative Policy Optimization (GRPO) against an untrained baseline on a Constraint Awareness Benchmark. Both fine‑tuning methods increase behavioral compliance from 4% to about 90% but reduce explicit constraint reporting and erode retained third‑person knowledge, with GRPO showing more destructive effects. The results suggest that reward‑based signals may suppress constraints context‑independently, and that models fail to enumerate constraints on request even when they can avoid them internally.

By Arin Agarwal
arXiv AI
Sep 12

SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics

SemVerBench is a benchmark that evaluates how well large language models (LLMs) understand and apply version-constraint resolution semantics, such as determining whether a version satisfies constraints like ^1.2.3 or >=2.0. The study finds that many models struggle with certain corner cases, with GPT‑5.1 performing poorly while Claude and Opus perform much better. The authors suggest that the failures stem from an activation/application gap rather than a lack of knowledge, and recommend that coding agents delegate version resolution to a dedicated resolver tool.

By Qibai Chen, Zeming Liu
arXiv AI
Aug 26

Constraint-Guided Enterprise Data Mapping with Large Language Models

The paper introduces Constraint‑Guided Enterprise Data Mapping (CGM), a neuro‑symbolic approach that uses schema‑grounded admissibility constraints to steer large language models (LLMs) in aligning enterprise data. CGM operates in three stages: defining constraints with metadata, generating candidates under relaxed constraints to ensure feasibility, and ranking them with a bounded LLM. Experiments show that hard constraints dramatically reduce candidate space and improve F1 scores, enabling small models to match or surpass large LLMs at a fraction of the cost while reducing expert effort.

By Sebastian Monka, Pramod Anantharam, Thien Vo Minh, Lavdim Halilaj
arXiv AI
Aug 14

Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues

arXiv:2608. 12599v1 Announce Type: new Abstract: Multi-turn dialogues let users revoke constraints as easily as impose them, but revocation does not reliably take effect: models keep enacting withdrawn requirements (occasionally beneath comments asserting their removal), a failure we call \emph{behavioral relapse}, or revocation inertia.

By Haoyuan Zhu
arXiv AI
Aug 11

Improving Constraint Models with LLM Agents

arXiv:2608. 08127v1 Announce Type: new Abstract: The runtime of Constraint Programming (CP) solvers is highly sensitive to modeling choices, such as symmetry breaking, implied constraints, global constraints, constraint reformulation, and variable representation.

By Florentina Voboril, Stefan Szeider
arXiv AI
Sep 4

It's the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories

The paper investigates whether large language models’ reasoning traces truly contain early, informative signals or merely reflect budget and difficulty confounds. Using a restart‑controlled truncation probe, the authors compare continuation success rates against from‑scratch restart curves across 178 problem‑model pairs, finding that only one case shows prefix‑limited success and that continuing a model’s own prefix generally outperforms restarting. A difficulty‑controlled test and two generation‑free analyses reveal that early internal signals do not carry outcome information beyond a problem‑difficulty baseline, underscoring the need for proper counterfactual controls.

By Yigit Utku Bulut