arXiv Machine Learning By Irena Girshovitz, Dan Zeltzer, Ran Gilad-Bachrach

Automated Synthesis and Adversarial Validation of Executable Causal Research Pipelines

Read the original on arXiv Machine Learning →

arXiv:2607. 21173v1 Announce Type: new Abstract: While automated research systems promise to accelerate empirical analysis, they are prone to silent failures: instances in which analysis code executes successfully yet relies on invalid causal assumptions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jul 23

Automated Synthesis and Adversarial Validation of Executable Causal Research Pipelines

While automated research systems promise to accelerate empirical analysis, they are prone to silent failures: instances in which analysis code executes successfully yet relies on invalid causal assumptions. We present the Artificial Intelligence (AI)-based Epidemiology Research Assistant (ARA), a framework that makes these failures visible by explicitly encoding causal design principles, study-specific assumptions, and methodological constraints.

arXiv Machine Learning
Sep 11

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

CausalArena is a new benchmark designed to evaluate causal discovery methods in the era of foundation models. It unifies synthetic structural causal models (SCMs), semantically grounded SCMs, and formula‑grounded SCMs, while also including real‑world datasets for external validation. Experiments show that performance rankings vary widely across different SCM families and protocols, indicating that strong results on one benchmark do not necessarily transfer to others.

By Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye
arXiv AI
Aug 20

CausalProfiler: Generating Synthetic Benchmarks for Rigorous and Transparent Evaluation of Causal Machine Learning

CausalProfiler is a synthetic benchmark generator designed to evaluate causal machine learning (Causal ML) methods more rigorously and transparently. It randomly samples causal models, data, queries, and ground truths based on explicit design choices across observation, intervention, and counterfactual reasoning levels, providing coverage guarantees and transparent assumptions. The authors demonstrate its utility by testing several state‑of‑the‑art methods under diverse conditions, both within and outside the identification regime, highlighting the insights CausalProfiler can reveal.

By Panayiotis Panayiotou, Audrey Poinsot, Alessandro Leite, Nicolas Chesneau, Marc Schoenauer, \"Ozg\"ur \c{S}im\c{s}ek
arXiv AI
Sep 2

Medical Causal Hypothesis Verification with Large Language Models

The paper "Medical Causal Hypothesis Verification with Large Language Models" reports a small-scale study evaluating eight LLMs on 17 medical causal hypotheses. The authors introduce an evaluation framework and annotate 1,067 evidence points across six criteria, using nine metrics to assess performance. Results show that while LLMs have strong recall, they frequently fail to provide valid scientific articles, evidence, or reject unsupported hypotheses, revealing a critical limitation for their use in healthcare.

By Safiyyah Ahmed, Abrar Ansari, Md Aminul Islam, Elena Zheleva