The paper reports that counterfactual fairness audits of clinical language‑model agents are unreliable without accounting for a per‑action instability floor. By repeatedly running identical vignettes, the authors found that actions changed 8.7% of the time, with instability varying eightfold across actions. A second model confirmed a pooled floor of 6.7%, showing that any reported fairness estimate lacking this floor cannot be interpreted as evidence of disparity.
By Rohith Reddy Bellibaltu, Manpreet Singh, Deepak Parashar, Rahul Joshi
arXiv:2609.03221v3 Announce Type: replace-cross
Abstract: Counterfactual fairness audits of clinical language-model agents report a flip rate: how often an action changes when only the patient's demo...
By Rohith Reddy Bellibatlu, Manpreet Singh, Deepak Parashar, Rahul Joshi
The paper proposes a scalable, automated method for auditing candidate‑job matching systems for demographic bias. It employs large‑language‑model agents to generate neutral resumes, injects controlled demographic variations, ranks candidates with a fine‑tuned embedding model, and evaluates nine fairness metrics across counterfactual, group‑fairness, and merit‑aware families, producing a composite risk report. Experiments on a small corpus show that single‑score audits miss nuanced issues, underscoring the need for multi‑metric evaluation and LLM‑generated audits as a low‑cost complement to human reviews.
By Sai Yashwant, Shruti Bansal, Anurag Dubey, Samaroha Chatterjee, Satyam Kumar, Shreyash Gupta, Gantala Thulsiram
The paper reports the first systematic audit of open‑weight large language models (LLMs) in hiring contexts, examining how job‑posting language influences recruiter and job‑seeker simulations across six models. It finds that agentic language lowers recruiter scores for female candidates while communal language mitigates this effect, and that coded‑exclusion language sharply reduces recruiter scores for non‑White candidates and discourages non‑White personas from applying. The study also identifies the explicit demographic label as the main causal factor and proposes a concrete pre‑deployment audit protocol aligned with EU and U.S. regulatory requirements.
By Kosuke Kitahara, Nobuhiro Yamaguchi
arXiv:2609.22090v1 Announce Type: new
Abstract: An LLM producing the response pattern associated with a human psychological effect is not the same claim as the LLM possessing that bias. We present Ps...
By Joy Bose
The study evaluates counterfactual bias in ten open‑source large language models (LLMs) for pediatric Emergency Severity Index (ESI) prediction. By creating paired clinical vignettes that differ only in demographic or socioeconomic variables, the authors measure shifts in acuity assignment, finding that counterfactual sensitivity varies widely across model families and sizes. A fine‑tuned Qwen2.5‑7B model exhibited the lowest sensitivity, while larger or medical‑domain models sometimes showed greater shifts, highlighting the need for fairness assessment before clinical deployment.
By Manar Aljohani, Brandon Ho, Kenneth McKinley, Dennis Ren, Xuan Wang
arXiv:2605. 12530v2 Announce Type: replace-cross Abstract: LLM fairness should be evaluated through in-situ behavioral pattern rather than standardized-test Q&A benchmarks.
By Zeyu Tang, Sang T. Truong, Deonna Owens, Shreyas Sharma, Yibo Jacky Zhang, Brando Miranda, Sanmi Koyejo
arXiv:2605. 03217v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in settings that require nuanced ethical reasoning, yet existing bias evaluations treat model outputs as simply "biased" or "unbiased.
By Yash Aggarwal, Atmika Gorti, Vinija Jain, Aman Chadha, Krishnaprasad Thirunarayan, Manas Gaur
arXiv:2605.01048v2 Announce Type: replace-cross
Abstract: Counterfactual prompting (i.e., perturbing a single factor and measuring output change) is widely used to evaluate things like LLM bias and C...
By Zihao Yang, Mosh Levy, Yoav Goldberg, Byron C. Wallace
arXiv:2607. 28934v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly involved in the distribution of scarce resources, raising concerns about biased allocations based on characteristics like race and gender.
By Martin Lukk (University of Toronto)
The study examines a two‑agent résumé screening process where both employer‑side and candidate‑side agents exchange evidence before deciding who advances, contrasting it with the traditional one‑call automated screening. Using GPT‑5.5 and Claude Opus 4.7 on 600 constructed résumé‑job pairs, the two‑agent method increased the proportion of applications advanced (up to 39.3% for GPT‑5.5) and raised pass rates for borderline cases from 4.5% to 26.2% (GPT‑5.5) and 6.5% to 16.1% (Opus 4.7). The results show that the screening procedure itself, rather than just the underlying model, determines which candidates reach human review and how consistently that access recurs.
By Jian Gao, Hang Jiang
arXiv:2606. 17165v1 Announce Type: cross Abstract: Organizations and researchers show increasing interest in using large language models (LLMs) in place of human participants in A/B tests, in the hope of experimenting faster and at lower cost.
By Joel Persson, M{\aa}rten Schultzberg, Sebastian Ankargren