arXiv Computation and Language

Constrained Decoding Eliminates Structural Failures in Small LLMs but Reveals a Scale-Dependent Semantic Gap

arXiv Computation and Language
Sep 1

How Order-Sensitive Are LLMs? OrderProbe for Deterministic Structural Reconstruction

arXiv:2601.08626v4 Announce Type: replace Abstract: Large language models (LLMs) excel at semantic understanding, yet their ability to reconstruct internal structure from scrambled inputs remains und...

By Zhaolu Kang, Yingjie He, Kehan Jiang, Leqi Zheng, Jiachen Qian, Qianyuan Zhang, Chunlei Meng, Yujie Feng, Yuan Wang, Stephen Dou, Aming Wu, Pengxiang Zhao, Jiaxin Liu, Guansu Wang, Zeyu Zhang, Lei Wang, Qishi Zhan, Xiaomin He, Meisheng Zhang, Jianyuan Ni, Richeng Xuan
arXiv AI
6d ago

Schema-Key Wording as an Instruction Channel in Structured Generation under Constrained Decoding

The paper investigates how the wording of schema‑key tokens can serve as an implicit instruction channel in constrained decoding for structured generation. By treating structured generation as a multi‑channel instruction problem, the authors analyze when an instructional advantage is preserved after grammar projection and conduct experiments on GSM8K and Math500 across seven language models. Results show that changing only the schema‑key wording can significantly alter accuracy, with both positive and negative effects, and that prompt‑level and schema‑level instructions interact non‑additively.

By Yifan Le
arXiv Machine Learning
Aug 31

Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks

The paper introduces CodeRQ-Bench, the first benchmark for assessing large language model reasoning quality across coding tasks such as generation, summarization, and classification. It analyzes over a thousand mismatches from existing evaluators, identifies recurring limitations, and derives design insights that lead to a new two‑stage evaluator, VERA. Experiments show VERA outperforms strong baselines, improving AUCROC by up to 0.26 and AUPRC by up to 0.21 on four datasets.

By Yuangang Li, Justin Tian Jin Chen, Ethan Yu, David Hong, Iftekhar Ahmed