arXiv Statistics ML

Mining Causality: AI-Assisted Search for Instrumental Variables

Hugging Face Trending Papers
Sep 10

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

CausalArena is a unified, evolvable benchmark designed to evaluate causal discovery methods across diverse structural causal models (SCMs). It incorporates synthetic SCMs for controlled structural variation, semantic operational SCMs for human-auditable environments, and formula-grounded SCMs to test discovery under explicit scientific mechanisms, along with real-world datasets for external validity. Experiments show that performance rankings vary significantly across SCM families and protocols, indicating that strong results on one benchmark do not generalize to others, especially in the context of causal discovery foundation models.

arXiv Machine Learning
Sep 11

AI Economist Agent: An Agentic Framework for Evidence-Based Economic and Financial Analysis with RAG, Knowledge Graphs, and Large Language Models

The paper introduces an AI economist agent that integrates large language models, retrieval‑augmented generation, knowledge graphs, and quantitative models to conduct evidence‑based economic and financial scenario analysis. The framework orchestrates LLM agents to plan analyses, retrieve relevant evidence, and structure economic mechanisms, while registered quantitative models produce numerical outcomes and predefined tests validate intermediate results for inclusion in the final report. Applied to European macro‑financial stress scenarios and bank capital analysis, the empirical study demonstrates the agent’s ability to combine flexible evidence retrieval and scenario construction while maintaining traceability to sources and explicit model calculations.

By Masahiro Kato
arXiv Machine Learning
Sep 11

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

CausalArena is a new benchmark designed to evaluate causal discovery methods in the era of foundation models. It unifies synthetic structural causal models (SCMs), semantically grounded SCMs, and formula‑grounded SCMs, while also including real‑world datasets for external validation. Experiments show that performance rankings vary widely across different SCM families and protocols, indicating that strong results on one benchmark do not necessarily transfer to others.

By Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye