arXiv Computation and Language By Bohan Yu, Shi-Yang Li, Pengfei Cao, Jun Zhao, Kang Liu

RuleWeaver: Benchmarking Rule-Centered Scenario Reasoning for Large Language Models

Read the original on arXiv Computation and Language →

RuleWeaver is a benchmark construction framework designed to evaluate large language models’ ability to reason over complex, rule‑centered scenarios. It begins with corpus‑derived IF‑THEN meta rules, expands them into more intricate rules, and composes these into scenario‑based QA instances. The benchmark assesses not only final answer correctness but also process‑level metrics such as rubric‑based answer quality, rule recall, and rule precision, revealing that current LLMs achieve only about 50% of the maximum rubric score on these tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Aug 25

Beyond Factual Knowledge: Benchmarking and Learning Step-Level Procedural Rule Reasoning in Large Language Models

arXiv:2608.22753v1 Announce Type: new Abstract: Large language models (LLMs) excel at text understanding and generation, yet still struggle to reliably understand and apply externally provided proced...

By Bohan Yu, Pengfei Cao, Chen Han, Chenxi Zhou, Zhiheng Zhang, Zhiyang Xie, Wenhao Teng, Xiangwen Liao, Jun Zhao, Kang Liu
arXiv Computation and Language
Sep 14

Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models

The paper introduces Tasks over Application Manuals (TAM), a benchmark designed to test long‑horizon procedural reasoning in large language models. TAM uses real‑world tasks from ICD‑10‑CM clinical coding and U.S. federal sentencing, requiring models to follow extensive, rule‑based manuals and perform interdependent steps to produce exact answers. Experiments with GPT‑5 and various prompting strategies show very low exact‑match accuracy—1% for coding and 15.5% for sentencing—highlighting a gap between current benchmarks and the ability to reliably follow complex procedures.

By Utkarsh Soni, Syed Shariyar Murtaza, Yifan Nie, Sachin Chandrasekhar, Eugene Wen
Hugging Face Trending Papers
5d ago

RGDT-Bench: Benchmarking LLM Reasoning for Rule-Governed Decisions and Their Justifications

RGDT-Bench is a new benchmark that evaluates large language models on Rule‑Governed Decision Tasks, where models must apply external rules to facts, justify decisions, and provide checkable justifications. The benchmark offers 202.1K condition‑level supervision slots across four task tracks and eight task‑probe combinations, and it labels warrant completeness through label‑blind extraction and deterministic checks. Evaluation shows that among correct responses, 40.2% of warrants are incomplete, and existing evaluators struggle to detect this, prompting the authors to train a reward model that improves AUROC to 69.24% and outperforms outcome‑supervised baselines.