arXiv Machine Learning By Jiale Zhang, Michael Larionov, Zichong Wang, Zhipeng Yin, Wenbin Zhang

FairLMs: A Turnkey Library for Fairness in Language Models

Read the original on arXiv Machine Learning →

FairLMs is a Python library designed to streamline fairness research in language models by unifying bias measurement, mitigation, and evaluation evidence. It offers 33 intrinsic and extrinsic metrics, 14 mitigation components across four intervention categories, 14 diagnostic tools, adapters for major Transformer architectures and hosted APIs, and benchmark loaders. The library enforces explicit declarations of model capabilities and input requirements, ensuring compatibility and reproducibility across components and datasets.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 17

FairCompressAgent: An Agentic Framework for Fairness-Aware Model Compression for FPGA Deployment

FairCompressAgent (FCA) is an agentic framework that unifies fairness-aware pruning, incremental quantization, and sparse low‑rank factorization for FPGA deployment. A language‑model planner selects compression configurations based on model profiles and measured outcomes, while an execution layer handles compression, fine‑tuning, evaluation, and constraint‑based selection. Experiments on Fitzpatrick‑17k with VGG‑11 show FCA can reduce inference tensor storage by 59.54% under accuracy constraints, improve validation average precision, and lower equalized opportunity, achieving similar results to one‑shot planning with fewer candidate evaluations.

By Yuanbo Guo, Yiyu Shi
arXiv AI
Aug 19

Position: Fairness Failure in Generative Models is an Evaluation Problem

The paper argues that fairness failures in generative models arise mainly from inadequate evaluation practices, making fairness findings hard to compare or use for deployment. It diagnoses common empirical and conceptual shortcomings in current methods and calls for a move toward standardized, generative‑specific evaluation. The authors introduce Fairness Cards, a minimal reporting artifact that explicitly documents evaluation choices—such as prompt families, counterfactual protocols, metrics, and refusal handling—to improve reproducibility, comparability, and accountability.

By Mariia Vladimirova, Jean-Yves Franceschi, Thibaut Issenhuth