arXiv AI By Joel Persson, M{\aa}rten Schultzberg, Sebastian Ankargren

Statistical Foundations of LLM-based A/B Testing: A Surrogacy Framework for Human Causal Inference

Read the original on arXiv AI →

arXiv:2606. 17165v1 Announce Type: cross Abstract: Organizations and researchers show increasing interest in using large language models (LLMs) in place of human participants in A/B tests, in the hope of experimenting faster and at lower cost.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 14

Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective

The paper discusses how large language models (LLMs) can be fine‑tuned with observational data to improve alignment with human preferences and business goals. It highlights that directly using such data can cause models to learn spurious correlations, and introduces DeconfoundLM, a method that removes known confounders from reward signals. Experiments show that DeconfoundLM better recovers causal relationships and outperforms baseline methods by over 16% in objective score when confounding is present.

By Erfan Loghmani