arXiv Computation and Language

RupeeBias: Auditing Demographic Bias in Indian Economic Guidance from Large Language Models

RupeeBias is a new benchmark that audits demographic bias in large language models (LLMs) when they provide economic guidance in India. It contains 39,150 prompts across four use cases—salary estimation, salary increment estimation, counter‑offer recommendation, and service pricing recommendation—varying 87 India‑specific demographic identifiers such as caste, religion, regional identity, gender, disability, and urban‑rural location. Evaluations of nine LLMs show that outputs differ by an average of 20.2% when only the demographic identifier changes, revealing systematic disparities across all six axes.

arXiv AI
Jun 17

Would a Large Language Model Pay Extra for a View? Inferring Willingness to Pay from Subjective Choices

arXiv:2602. 09802v2 Announce Type: replace Abstract: As Large Language Models (LLMs) are increasingly deployed in applications such as travel assistance and purchasing support, they are often required to make subjective choices on behalf of users in settings where no objectively correct answer exists.

By Manon Reusens, Sofie Goethals, Toon Calders, David Martens
arXiv AI
Sep 24

Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms

The paper introduces a generation benchmark for culturally specific kinship terms, evaluating five open‑weight large language models (LLMs) on Hindi, Tamil, and Korean. Unlike prior multiple‑choice tests that treat kinship understanding as a recognition task, the study prompts LLMs to generate terms across two communicative tasks and compares results to a matched option‑supported baseline. Findings show that while models like GPT‑OSS120B and Llama‑3.370B can select correct terms in over 90% and 78% of cases respectively, they produce the correct term only 36% and 24% of the time, indicating a significant evaluation‑format gap and highlighting the difficulty of culturally specific kinship generation even when relationships are explicitly stated.

By Sahil Pardasani, Madhusudan Singh
arXiv AI
Sep 16

Beyond the Name: Demographic Leakage in De-Identified R\'esum\'es and Evaluation Artifacts in LLM Bias Audits

The paper examines whether removing declared language fields from de‑identified résumés eliminates demographic leakage in large language models. By keeping language attributes identical and varying only unstructured prose across five ethnocultural groups and three cue‑salience levels, the authors find that non‑language text still allows target‑group recovery (average 0.757, reaching 1.000 under high salience). They also show that evaluation design—such as allowing or forbidding ties—dramatically affects LLM‑as‑a‑judge outcomes, underscoring the importance of evaluation protocol in bias audits.

By Qiangju Chen, Yang Xiao
arXiv AI
Jun 2

IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesian Languages

arXiv:2606. 01260v1 Announce Type: cross Abstract: Despite being home to more than 1300 ethnic groups and 700 indigenous languages, bias in Large Language Models has not been fully studied in Indonesia, thus leaving a critical gap in evaluating representational fairness and localized stereotypes within its uniquely vast, multilingual, and diverse sociocultural landscape.

By Ikhlasul Akmal Hanif, Muhammad Falensi Azmi, Filbert Aurelian Tjiaranata, Eryawan Presma Yulianrifat, Fajri Koto
arXiv Machine Learning
Jul 21

STRATA: A Name-and-Geography Race Inference Model for Fair Lending and Housing Equity Applications

arXiv:2504. 21259v2 Announce Type: replace-cross Abstract: Accurate imputation of race and ethnicity (R&E) is essential for fair lending compliance under ECOA, HMDA, and the Community Reinvestment Act, where up to 15% of mortgage applications carry missing race data and regulated institutions bear responsibility for identifying disparities on those records.

By S. Chalavadi, A. Pastor, T. Leitch