← Back to all news
arXiv Machine Learning August 24, 2026 By Zihao Yang, Mosh Levy, Yoav Goldberg, Byron C. Wallace

Compared to What? Baselines and Metrics for Counterfactual Prompting

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • llms
  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jun 17

Statistical Foundations of LLM-based A/B Testing: A Surrogacy Framework for Human Causal Inference

arXiv:2606. 17165v1 Announce Type: cross Abstract: Organizations and researchers show increasing interest in using large language models (LLMs) in place of human participants in A/B tests, in the hope of experimenting faster and at lower cost.

By Joel Persson, M{\aa}rten Schultzberg, Sebastian Ankargren
llmssafety
More like this →
arXiv AI
Aug 18

Demographic Injection in Medical Language Models under Diversity, Equity, and Inclusion Prompts

arXiv:2608. 15254v1 Announce Type: new Abstract: Clinical-AI guidance increasingly recommends prompting language models to reason with attention to diversity, equity, and inclusion (DEI).

By Diego Mardian, Frank Liu
llmsbenchmarks
More like this →
arXiv AI
Jun 16

AgentFairBench: Do LLM Agents Discriminate When They Act?

arXiv:2606. 16723v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly take actions (screening applicants, recommending credit, triaging patients), yet fairness for LLMs is still measured by grading answers.

By Triveni Morla, Rohith Reddy Bellibaltu, Manpreet Singh, Manmeet Singh Kapoor
llmsagentsbenchmarkssafety
More like this →
arXiv AI
Aug 11

See Me, Believe Me: Causality, Intersectionality, and Interventions Improving the Appearance of Patients

arXiv:2410. 01227v2 Announce Type: replace-cross Abstract: In the context of medical records, patients often experience testimonial injustice, where the textual account undermines the validity of their experiences.

By Kenya S. Andrews, Mesrob I. Ohannessian, Elena Zheleva
llms
More like this →
arXiv AI
Jun 8

When Large Language Models Fail in Healthcare: Evaluating Sensitivity to Prompt Variations

arXiv:2606. 07237v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in healthcare for tasks such as clinical question answering, diagnosis support, and report summarization.

By Mahdi Alkaeed
llmsnlpbenchmarkssafety
More like this →
arXiv AI
Aug 18

Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents

arXiv:2608. 14606v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whether answers look plausible at the individual level.

By Mantas Lukauskas, Viktorija \v{S}arkauskait\.e
llmsbenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea