CausalGame: Benchmarking Causal Thinking of LLM Agents in Games
arXiv:2607. 04293v1 Announce Type: cross Abstract: Building AI Scientist agents with Large Language Models (LLMs) has recently attracted growing attention.
arXiv:2607. 04293v1 Announce Type: cross Abstract: Building AI Scientist agents with Large Language Models (LLMs) has recently attracted growing attention.
arXiv:2606. 10607v1 Announce Type: cross Abstract: Causal discovery aims to uncover causal structures from observational data, which is crucial for real-world decision-making.
arXiv:2602. 16481v2 Announce Type: replace Abstract: Causal discovery seeks to uncover causal relations from data, typically represented as causal graphs, and is essential for predicting the effects of interventions.
arXiv:2606. 07525v1 Announce Type: cross Abstract: Causal graphs in text are typically populated by observable, predefined events.
Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use. Yet the relevant benchmark landscape largely divides into symbolic causal reasoning benchmarks without realistic data analysis or data analysis benchmarks without a principled causal data-generating structure.
arXiv:2609.39406v1 Announce Type: new Abstract: Causal AI is a branch of Artificial Intelligence which helps understand and reason about cause and effect relationships, not just patterns or correlati...
arXiv:2607. 08093v1 Announce Type: new Abstract: Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use.
arXiv:2404.06349v3 Announce Type: replace Abstract: The ability to understand causality significantly impacts the competence of large language models (LLMs) in output explanation and counterfactual r...
CausalArena is a unified, evolvable benchmark designed to evaluate causal discovery methods across diverse structural causal models (SCMs). It incorporates synthetic SCMs for controlled structural variation, semantic operational SCMs for human-auditable environments, and formula-grounded SCMs to test discovery under explicit scientific mechanisms, along with real-world datasets for external validity. Experiments show that performance rankings vary significantly across SCM families and protocols, indicating that strong results on one benchmark do not generalize to others, especially in the context of causal discovery foundation models.
The paper introduces an AI economist agent that integrates large language models, retrieval‑augmented generation, knowledge graphs, and quantitative models to conduct evidence‑based economic and financial scenario analysis. The framework orchestrates LLM agents to plan analyses, retrieve relevant evidence, and structure economic mechanisms, while registered quantitative models produce numerical outcomes and predefined tests validate intermediate results for inclusion in the final report. Applied to European macro‑financial stress scenarios and bank capital analysis, the empirical study demonstrates the agent’s ability to combine flexible evidence retrieval and scenario construction while maintaining traceability to sources and explicit model calculations.
arXiv:2607. 15281v1 Announce Type: new Abstract: Causal and intervention-based question answering is fundamental to advancing large language models (LLMs) toward reasoning beyond surface-level correlations and understanding underlying causal mechanisms.
CausalArena is a new benchmark designed to evaluate causal discovery methods in the era of foundation models. It unifies synthetic structural causal models (SCMs), semantically grounded SCMs, and formula‑grounded SCMs, while also including real‑world datasets for external validation. Experiments show that performance rankings vary widely across different SCM families and protocols, indicating that strong results on one benchmark do not necessarily transfer to others.