arXiv Computation and Language

Correct Prediction, Wrong Steps? Consensus Reasoning Knowledge Graph for Robust Chain-of-Thought Synthesis

arXiv Computation and Language
Sep 21

When Does Reasoning Help in Machine Translation? A Hierarchical Analysis of LRM Reasoning Traces

The paper investigates when intermediate reasoning traces benefit machine translation by examining models, languages, domains, and datasets. It finds that the optimal reasoning language depends on the model, reasoning length has a non‑monotonic effect on quality, and traces display recurring functional patterns. Using Hierarchical Meta‑Summarization, the authors uncover a shared structure of understanding/planning, translating/drafting, and refining/verifying, while also noting domain‑specific variations, suggesting that reasoning should be tailored rather than uniformly applied.

By Yuxiang Liu, Jiaming Luo, Eleftheria Briakou, Colin Cherry
arXiv Computation and Language
Sep 1

Reasoning over Grammar: Can Synthetic Linguistic Reasoning Traces Enhance Low-Resource Machine Translation?

The paper explores whether structured linguistic reasoning traces can improve low‑resource machine translation by guiding large language models (LLMs). It proposes a pipeline that automatically generates step‑by‑step reasoning traces from Universal Dependencies treebanks, dictionaries, and grammar‑rule banks, and evaluates these traces in in‑context learning, supervised fine‑tuning, and reinforcement fine‑tuning on Xibe and Chintang. The results show that providing reliable reasoning traces at inference time significantly boosts translation quality, whereas using them as training data yields smaller, less consistent gains, indicating that LLMs can benefit from grammatical guidance but struggle to generate accurate analyses themselves.

By Renhao Pei, Yihong Liu, Sampo Pyysalo, Hinrich Sch\"utze, Shaoxiong Ji
arXiv Computation and Language
Aug 28

TRACES: Tagging Reasoning Steps for Adaptive Cost-Efficient Early-Stopping

TRACES (Tagging Reasoning Steps for Adaptive Cost‑Efficient Early‑Stopping) is a lightweight framework that tags reasoning steps of large‑language models in real time, enabling adaptive, cost‑efficient early stopping during inference. By monitoring the types of steps generated, the method identifies when models shift their reasoning after arriving at a correct answer, allowing for interpretable stopping criteria. Experiments on mathematical reasoning benchmarks (MATH500, GSM8K, AIME) and knowledge benchmarks (MMLU, GPQA) show token reductions of 20–50% while preserving accuracy, with more conservative thresholds needed for harder tasks such as BeyondAIME and IMO AnswerBench.

By Yannis Belkhiter, Seshu Tirupathi, Giulio Zizzo, John D. Kelleher
arXiv AI
Aug 28

Improving LLM Interpretability with User-Centric Chain-of-Thought Reasoning

The paper proposes a user‑centric Chain‑of‑Thought (CoT) reasoning framework that structures LLM reasoning traces into self‑contained, verifiable steps using XML‑like tags. This design allows users to independently assess and correct the AI’s reasoning while preserving performance on mathematical reasoning tasks. User studies show that the approach improves perceived usefulness and ease of use compared to standard CoT.

By Philipp Schr\"oppel
arXiv Computation and Language
Aug 27

DCGC: Draft-Conditioned Global Correction for Complex Reasoning with Masked Diffusion Models

DCGC is a Masked Diffusion Model framework that performs global correction of flawed reasoning traces in Large Language Models. It uses an imperfect solution draft from an upstream solver as auxiliary context and combines task‑specific supervised fine‑tuning with a Dynamic Dual‑CFG inference mechanism that separates problem‑only and joint problem‑draft branches. Experiments on math, code, and knowledge reasoning benchmarks show that DCGC outperforms standard sampling and simpler CFG variants, and it can improve full test‑set accuracy even when ground‑truth failure labels are unavailable.

By Minhae Oh, Nakyung Lee, Jungwoo Lee