Learning to reason with LLMs
Related stories
SmolLM3: smol, multilingual, long-context reasoner
Controlling Reasoning Effort in LLMs
How LLMs Learn Low-, Medium-, and High-Effort Reasoning Modes
How reliable are LLMs when it comes to playing dice?
arXiv:2606. 07515v1 Announce Type: cross Abstract: We investigate the probabilistic reasoning capabilities of large language models through a controlled benchmarking study on discrete probability problems.
Categories of Inference-Time Scaling for Improved LLM Reasoning
And an Overview of Recent Inference-Scaling Papers
Can LLMs Reason Structurally? Benchmarking via the Lens of Data Structures
arXiv:2505. 24069v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are deployed on increasingly complex tasks that require multi-step decision-making.
Structure-Internalized Rule Language Model for Faithful Knowledge Graph Reasoning
arXiv:2608. 17443v1 Announce Type: new Abstract: Knowledge Graph Reasoning (KGR) aims to discover latent facts by leveraging the structural evidence available in KGs, posing a challenge to the structural semantic understanding capability of KGR models.
Can Post-Training Transform LLMs into Causal Reasoners?
arXiv:2602. 06337v2 Announce Type: replace-cross Abstract: Causal inference is essential for decision-making but remains challenging for non-experts.
Project Auto-World: Towards Automated Benchmarking of Neural Relational Reasoners
arXiv:2606. 24965v1 Announce Type: cross Abstract: Reasoning about relational structures remains a significant challenge for neural models, particularly when they must systematically apply learned knowledge to problem instances that are harder than those seen in training.
X-RAY: Mapping LLM Reasoning Capability via Formalized and Calibrated Probes
arXiv:2603. 05290v2 Announce Type: replace Abstract: Large language models (LLMs) achieve promising performance, yet their ability to reason remains poorly understood.
HOLMES: Evaluating Higher-Order Logical Reasoning in LLMs
arXiv:2606. 23238v2 Announce Type: replace Abstract: Logical reasoning is essential for reliable AI, yet existing benchmarks are largely first-order-logic-centric, focusing on object-level deduction over fixed predicates.
Enhancing LLMs' Clinical Reasoning with Real-World Data from a Nationwide Sepsis Registry
arXiv:2505. 02722v2 Announce Type: replace Abstract: Although large language models (LLMs) have demonstrated impressive reasoning capabilities across general domains, their effectiveness in real-world clinical practice remains limited.

