arXiv AI By Eliza Berman, Bella Chang, Daniel B. Neill, Emily Black

Attribution Bias in Large Language Models

Read the original on arXiv AI →

The paper introduces AttriBench, a benchmark dataset that balances author fame and demographics to study quote attribution in large language models (LLMs). Using AttriBench, the authors evaluate 11 popular LLMs and find that accurate attribution remains difficult, with significant disparities across race, gender, and intersectional groups. They also identify a new failure mode—suppression—where models omit attribution entirely, which is unevenly distributed across demographics and not reflected by standard accuracy metrics.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 3

Probing Cultural Signals in Large Language Models through Author Profiling

The study investigates cultural biases in large language models (LLMs) by testing their ability to perform author profiling—inferring singers’ gender and ethnicity—from song lyrics in a zero‑shot setting. Evaluating over 10,000 lyrics across several open‑source models, the authors find that most LLMs default toward North American ethnicity, while DeepSeek‑1.5B leans toward Asian ethnicity, and that Ministral‑8B exhibits the strongest ethnicity bias whereas Gemma‑12B is the most balanced. The paper introduces two fairness metrics, Modality Accuracy Divergence (MAD) and Recall Divergence (RD), to quantify these disparities and provides code and results publicly on GitHub and HuggingFace.

By Valentin Lafargue, Ariel Guerra-Adames, Emmanuelle Claeys, Elouan Vuichard, Jean-Michel Loubes
arXiv Computation and Language
Aug 21

The Asymmetric Harms of LLM Compression

arXiv:2608. 19670v1 Announce Type: new Abstract: Large language models (LLMs) compression reduces deployment costs, but standard aggregate metrics like perplexity and accuracy often mask underlying behavioral shifts.

By Yuan Wu, Mairui Li, Lesia Semenova, Chudi Zhong
arXiv AI
Sep 10

Attribution in Scientific Literature: New Benchmark and Methods

The paper introduces REASONS, a benchmark of 12,723 sentence-level citation instances across 12 arXiv subject categories, to evaluate scientific citation attribution under different evidence conditions. It proposes a dual-metric framework—Abstention Rate (AR) and Hallucination Rate (HR)—to balance reliability and responsiveness. Experiments with proprietary and open-source LLMs across various prompting and retrieval settings show that advanced Retrieval-Augmented Generation (RAG) reduces hallucinations but increases abstention, while adversarial metadata can push hallucination rates above 85%. Human evaluation confirms a high ratio of factual hallucinations to acceptable paraphrases, underscoring the need for systems that can appropriately abstain under uncertainty.

By Deepa Tilwani, Yash Saxena, Seyedali Mohammadi, Ankur Padia, Edward Raff, Amit Sheth, Srinivasan Parthasarathy, Manas Gaur