arXiv Computation and Language By Pavithra P M Nair, Bhavik Talaviya, Shourya Bhushan, Rahul Pankajakshan, Seema Guruvadoo, Avinash Agarwal, Gilad Gressel, Krishnashree Achuthan

RupeeBias: Auditing Demographic Bias in Indian Economic Guidance from Large Language Models

Read the original on arXiv Computation and Language →

RupeeBias is a new benchmark that audits demographic bias in large language models (LLMs) when they provide economic guidance in India. It contains 39,150 prompts across four use cases—salary estimation, salary increment estimation, counter‑offer recommendation, and service pricing recommendation—varying 87 India‑specific demographic identifiers such as caste, religion, regional identity, gender, disability, and urban‑rural location. Evaluations of nine LLMs show that outputs differ by an average of 20.2% when only the demographic identifier changes, revealing systematic disparities across all six axes.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jun 17

Would a Large Language Model Pay Extra for a View? Inferring Willingness to Pay from Subjective Choices

arXiv:2602. 09802v2 Announce Type: replace Abstract: As Large Language Models (LLMs) are increasingly deployed in applications such as travel assistance and purchasing support, they are often required to make subjective choices on behalf of users in settings where no objectively correct answer exists.

By Manon Reusens, Sofie Goethals, Toon Calders, David Martens
arXiv AI
Sep 24

Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms

The paper introduces a generation benchmark for culturally specific kinship terms, evaluating five open‑weight large language models (LLMs) on Hindi, Tamil, and Korean. Unlike prior multiple‑choice tests that treat kinship understanding as a recognition task, the study prompts LLMs to generate terms across two communicative tasks and compares results to a matched option‑supported baseline. Findings show that while models like GPT‑OSS120B and Llama‑3.370B can select correct terms in over 90% and 78% of cases respectively, they produce the correct term only 36% and 24% of the time, indicating a significant evaluation‑format gap and highlighting the difficulty of culturally specific kinship generation even when relationships are explicitly stated.

By Sahil Pardasani, Madhusudan Singh