TempoBench: Evaluating Temporal Causal Reasoning in Large Language Models
arXiv:2510. 27544v2 Announce Type: replace Abstract: Temporal reasoning involves understanding how systems evolve over time through input-driven state transitions.
arXiv:2510. 27544v3 Announce Type: replace Abstract: Current training paradigms, optimized for long-horizon reasoning trace execution, have made Large Language Models (LLMs) excel at pattern matching and forward simulation of reasoning, but underperform at counterfactual causal understanding and reasoning.
arXiv:2510. 27544v2 Announce Type: replace Abstract: Temporal reasoning involves understanding how systems evolve over time through input-driven state transitions.
arXiv:2607. 08093v1 Announce Type: new Abstract: Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use.
arXiv:2602. 20094v2 Announce Type: replace Abstract: As large language models (LLMs) witness increasing deployment in complex, high-stakes decision-making scenarios, it becomes imperative to ground their reasoning in causality rather than spurious correlations.
Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use. Yet the relevant benchmark landscape largely divides into symbolic causal reasoning benchmarks without realistic data analysis or data analysis benchmarks without a principled causal data-generating structure.
arXiv:2606. 05402v1 Announce Type: cross Abstract: Large reasoning models (LRMs) produce reasoning traces with non-linear structures, such as backtracking and self-correction, that complicate the evaluation and monitoring of the reasoning process.
arXiv:2607. 18100v1 Announce Type: new Abstract: Extended reasoning has become standard for frontier Large Language Models (LLMs), yet the trajectories these models produce remain largely uncontrollable.
arXiv:2609.38409v1 Announce Type: new Abstract: Recent progress in large language model reasoning has been driven by benchmarks and reinforcement learning environments with automatically verifiable r...
The paper introduces Causal Shortcut Learning (CSL), a framework that identifies token chains—called causal shortcuts—that guide Diffusion Language Models (DLMs) toward correct reasoning paths. By extracting these shortcuts and applying parallel prioritized masking during training, CSL improves both convergence speed and generation accuracy. Experiments on several reasoning benchmarks and two base models show CSL outperforms existing SFT-variant baselines, achieving an average 1.92% improvement over SFT-only models and up to 4.20% on MATH-500.
TRACES (Tagging Reasoning Steps for Adaptive Cost‑Efficient Early‑Stopping) is a lightweight framework that tags reasoning steps of large‑language models in real time, enabling adaptive, cost‑efficient early stopping during inference. By monitoring the types of steps generated, the method identifies when models shift their reasoning after arriving at a correct answer, allowing for interpretable stopping criteria. Experiments on mathematical reasoning benchmarks (MATH500, GSM8K, AIME) and knowledge benchmarks (MMLU, GPQA) show token reductions of 20–50% while preserving accuracy, with more conservative thresholds needed for harder tasks such as BeyondAIME and IMO AnswerBench.
arXiv:2404.06349v3 Announce Type: replace Abstract: The ability to understand causality significantly impacts the competence of large language models (LLMs) in output explanation and counterfactual r...
arXiv:2605. 12519v2 Announce Type: replace-cross Abstract: Training language models to produce both correct answers and sound reasoning remains an open challenge.
arXiv:2602. 06337v2 Announce Type: replace-cross Abstract: Causal inference is essential for decision-making but remains challenging for non-experts.