CausalDS: Benchmarking Causal Reasoning in Data-Science Agents
arXiv:2607. 08093v1 Announce Type: new Abstract: Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use.
arXiv:2607. 04293v1 Announce Type: cross Abstract: Building AI Scientist agents with Large Language Models (LLMs) has recently attracted growing attention.
arXiv:2607. 08093v1 Announce Type: new Abstract: Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use.
Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use. Yet the relevant benchmark landscape largely divides into symbolic causal reasoning benchmarks without realistic data analysis or data analysis benchmarks without a principled causal data-generating structure.
arXiv:2510. 27544v3 Announce Type: replace Abstract: Current training paradigms, optimized for long-horizon reasoning trace execution, have made Large Language Models (LLMs) excel at pattern matching and forward simulation of reasoning, but underperform at counterfactual causal understanding and reasoning.
arXiv:2409.14202v4 Announce Type: replace-cross Abstract: The instrumental variables (IVs) method is a leading empirical strategy for causal inference. Finding IVs is a heuristic and creative process...
arXiv:2602. 02905v2 Announce Type: replace Abstract: Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating their capacity for verifiable discovery remains a central challenge.
arXiv:2404.06349v3 Announce Type: replace Abstract: The ability to understand causality significantly impacts the competence of large language models (LLMs) in output explanation and counterfactual r...
arXiv:2609.01526v1 Announce Type: new Abstract: Scientific agents must learn not only how to reason, but also what to believe. However, existing LLM agents typically express scientific hypotheses in...
arXiv:2605.26087v2 Announce Type: replace-cross Abstract: Frontier LLMs now perform strongly across a wide range of physics evaluations, but it is hard to disentangle genuine reasoning from recall of...
CausalArena is a new benchmark designed to evaluate causal discovery methods in the era of foundation models. It unifies synthetic structural causal models (SCMs), semantically grounded SCMs, and formula‑grounded SCMs, while also including real‑world datasets for external validation. Experiments show that performance rankings vary widely across different SCM families and protocols, indicating that strong results on one benchmark do not necessarily transfer to others.
arXiv:2606. 10607v1 Announce Type: cross Abstract: Causal discovery aims to uncover causal structures from observational data, which is crucial for real-world decision-making.
arXiv:2602. 06337v2 Announce Type: replace-cross Abstract: Causal inference is essential for decision-making but remains challenging for non-experts.
CausalArena is a unified, evolvable benchmark designed to evaluate causal discovery methods across diverse structural causal models (SCMs). It incorporates synthetic SCMs for controlled structural variation, semantic operational SCMs for human-auditable environments, and formula-grounded SCMs to test discovery under explicit scientific mechanisms, along with real-world datasets for external validity. Experiments show that performance rankings vary significantly across SCM families and protocols, indicating that strong results on one benchmark do not generalize to others, especially in the context of causal discovery foundation models.