arXiv:2609.24228v1 Announce Type: new
Abstract: Text-to-image (T2I) models are typically evaluated for bias using slot-based templates such as ``a photo of a [profession]''. Such templates probe only...
By Yue Dai, Ziyang Liu, Marc Cheong, Caren Han
arXiv:2508. 03483v3 Announce Type: replace-cross Abstract: While prior research on text-to-image generation has predominantly focused on biases in human depictions, demographic bias in generated objects remains relatively underexplored.
By Dasol Choi, Jihwan Lee, Minjae Lee, Minsuk Kahng
arXiv:2512. 00807v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) inherit significant social biases from their training data, notably in gender representation.
By Yujie Lin, Jiayao Ma, Qingguo Hu, Wenbo Li, Genji Li, Derek Wong, Jinsong Su
arXiv:2608.29590v1 Announce Type: new
Abstract: We propose a societal bias evaluation method for large vision-language models (LVLMs) in the era of strong safety guardrails. Existing benchmarks rely...
By Yusuke Hirota, Michael Ross Boone, Arun George Zachariah, Jibin Rajan Varghese, Yu-Chiang Frank Wang, Boyi Li, Ryo Hachiuma
arXiv:2609.16366v1 Announce Type: cross
Abstract: When foundation models describe people, recent work in AI fairness, accessibility, and ethics recommends avoiding inferred identity labels (e.g., "sh...
By Yingjia Wan, Lin Lin, Elisa Kreiss
The study introduces GAPA, a dataset of 316 physical attributes with 14,706 gender-association ratings from 304 US annotators, showing that such descriptions carry structured gender associations. It evaluates 16 LLMs, finding they partially mirror human ratings but exhibit biases such as compressed distributions, weaker alignment for men, and asymmetric abstention toward non‑binary identities. A proxy model trained on these data is released and applied to analyze character descriptions in LitBank, illustrating the persistence of gendered interpretations in ostensibly neutral language.
By Yingjia Wan, Lin Lin, Elisa Kreiss
arXiv:2601. 04946v3 Announce Type: replace-cross Abstract: Automatic metrics are widely used to evaluate text-to-image models, often replacing human judgment in benchmarking, model selection, and large-scale data filtering.
By Subhadeep Roy, Gagan Bhatia, Steffen Eger
Public trust in Autonomous Vehicles (AVs) may depend not only on technical success but also on the fairness of their decision making. While a recent trend in AV research involves using general purpose...
arXiv:2609.05540v1 Announce Type: cross
Abstract: Many medical conditions require diagnosis through detailed, multi-context clinical assessment rather than from visual appearance alone. Despite this,...
By Karan Dua, Amit Agarwal, Hitesh Laxmichand Patel, Hansa Meghwani, Jyotika Singh, Ranjeet Gupta, Graham Horwood, Tao Sheng, Avi Sil, Sujith Ravi, Dan Roth
arXiv:2608.21415v1 Announce Type: cross
Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable performance across a wide range of tasks; however, they often inherit social biases fro...
By Yisong Xiao, Aishan Liu, Yongxin Huang, Zonghao Ying, Shiji Zhao, Tianlin Li, Yong Han, Jian Yang, Xianglong Liu
FairLens is a benchmark and evaluation framework that measures fairness and validity of vision‑language models (VLMs) in high‑stakes domains such as hiring, legal, and healthcare. It uses over 100,000 face‑image and question pairs covering gender, race, and age, and assesses responses through demographic parity, soundness, demographic association, and bias in free‑text generation. The study finds that VLMs often make unwarranted inferences from faces rather than abstaining, especially in legal and healthcare contexts, and that small parity gaps can still hide unsafe treatment across groups.
By Vahid Reza Khazaie, Ahmed Y. Radwan, Shaina Raza
arXiv:2512. 04981v2 Announce Type: replace-cross Abstract: Text-to-image (T2I) systems increasingly rely on Large Language Model (LLM)-based text conditioning to interpret and expand user prompts.
By NaHyeon Park, Na Min An, Kunhee Kim, Soyeon Yoon, Jiahao Huo, Hyunjung Shim