arXiv AI By Naihao Deng, Samee Arif, Shuaichen Chang, Yulong Chen, Rada Mihalcea

One Example Is Enough to Pass Fairness Benchmarks: Rethinking Fairness Evaluation for Aligned LLMs

Read the original on arXiv AI →

The paper demonstrates that a single example from the BBQ fairness benchmark can dramatically improve a model’s performance, raising accuracy from 79.9% to 92.9% with Group Relative Policy Optimization and to 99.0% with one-shot in-context learning. This effect is consistent across different model families and is driven by the model’s reasoning traces, which adopt a category‑agnostic "missing evidence" pattern. The authors argue that BBQ-style multiple‑choice abstention tests capture only a single structural cue and therefore do not guarantee true fairness, calling for broader evaluation suites.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 10

Let's Unlearn Stereotypes Before Decision-Making: Assessing the Impact of Intrinsic Bias Mitigation on Downstream Fairness in LLMs

arXiv:2509. 16462v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly used in high-stakes decision-making systems, where biased predictions can reinforce social and economic disparities.

By Mina Arzaghi, Alireza Dehghanpour Farashah, Florian Carichon, Jean-Fran\c{c}ois Plante, Golnoosh Farnadi
arXiv AI
Jun 10

Is Fairness Truly Fair? Towards Reliable Lipschitz Fairness in Multi-Task Learning via Fixed-\texorpdfstring{$\delta$}{delta} Alignment

arXiv:2606. 10632v1 Announce Type: cross Abstract: Lipschitz-style individual fairness formalizes the idea that semantically similar examples should receive similar predictions, but its evaluation in multi-task learning (MTL) can be confounded by method-induced representation scales.

By Junbo Ding, Xin Zang, Chenchen Pan, Donghao Song, Jiaxin Zhu, Danhuai Guo
arXiv AI
Aug 17

Training Fair Tabular Foundation Models

arXiv:2608. 14211v1 Announce Type: cross Abstract: Tabular Foundation Models (TFMs) have emerged as leading methods for tabular predictive tasks, leveraging in-context learning to predict on new data without task-specific training.

By Patrik Kenfack, Jesse C. Cresswell, Anthony L. Caterini, Samira Ebrahimi Kahou, Ulrich A\"ivodji