arXiv Machine Learning

MIRAGE: Auditing Anti-Muslim Bias in Frontier LLMs Across Reasoning, Agentic, and Time-Coupled Conditions

arXiv:2606. 16562v1 Announce Type: new Abstract: Five years after the discovery of persistent anti-Muslim bias in large language models, most evaluations remain confined to single-turn prompt completion, a setting that no longer reflects how frontier LLMs are deployed.

arXiv AI
Aug 24

Ansari: A Retrieval-Grounded Islamic AI Assistant -- Architecture, Deployment, and Lessons from 140,000 Conversations

Ansari is a retrieval‑grounded Islamic AI assistant that has handled over 140,000 conversations in more than 25 languages since June 2023. It uses an agentic retrieval loop where a language model searches authenticated Islamic corpora—including the Qur’an, hadith collections, fiqh encyclopedias, and tafsir sources—and answers only based on retrieved content, providing citations for verification. The paper details Ansari’s architecture, multi‑platform deployment, evaluation results (including top performance on the IslamicMMLU leaderboard and strong resistance to false premises), and lessons for faith‑sensitive LLM deployments.

By M Waleed Kadous, Amr Elsayed, Abdullah Al Nahas, Ashraf Haress
arXiv Computation and Language
Sep 2

Evaluating Second-Order Bias of LLMs Through Epistemic Entitlement

The paper introduces a novel framework for assessing second‑order bias in large language models (LLMs), defined as bias in how an LLM judges the acceptability of biased content. Using principles from entitlement epistemology, the authors design a reasoning task that asks LLMs to determine whether a biased text is acceptable for specific demographic groups, and propose two metrics to quantify biased judgments. Experiments on both open‑source and closed‑source models reveal that the task bypasses safety guardrails, uncovers systematic variations across target groups, and demonstrates that models still rely on demographic labels when evaluating bias.

By Ramaravind Kommiya Mothilal, Terry Jingchen Zhang, Raiyan Ahmed, Zhijing Jin, Shion Guha, Syed Ishtiaque Ahmed
arXiv Computation and Language
Sep 18

Before the Arrest: Benchmarking LLMs on Criminal Profiling from Incomplete Evidence

The paper introduces the Profiling, Investigation, and Judgment (PIJ) benchmark, which contains 2,500 real homicide cases from five countries to evaluate large language models (LLMs) on pre‑arrest criminal investigation tasks. It assesses LLMs across criminal profiling, crime process reconstruction, and sentence prediction, revealing that performance drops as tasks require more implicit reasoning about unknown suspect profiles. The study finds that LLMs lag behind human experts, especially on inferential tasks like motivation and victim‑offender relationships, and exhibit biases in gender, age, and motive attribution.

By Yutong Yao, Yanjie Cao, Guanhua Chen, Xu Yang, Junchao Wu, Zeyu Wu, Lidia S. Chao, Derek F. Wong
arXiv AI
Sep 7

MABPD: Multi-Agent Bias Probing & Detection via Structured Argument Debate

MABPD (Multi‑Agent Bias Probing & Detection) is a training‑free pipeline that uses three specialized large language model agents to analyze news articles from complementary perspectives and resolve disagreements via a Structured Argument Debate (SAD) protocol. SAD imposes an asymmetric burden of proof—biased claims lacking grounded textual evidence receive zero weight—along with role‑weighted voting and post‑consensus verification, replacing task‑specific supervised decision boundaries. Ablation studies show that the debate module alone accounts for up to a 10.6‑point F1 gain, and on the BABE benchmark MABPD attains 83.4% macro F1, within 0.7 percentage points of the supervised state‑of‑the‑art, while achieving 75.0% zero‑shot accuracy on the SemEval 2019 HyperPartisan corpus.

By Garvit Joshi (Graphic Era University, Dehradun, India), Stavya Dhyani (Graphic Era University, Dehradun, India), Jasmine (Graphic Era University, Dehradun, India), Arun Chauhan (Graphic Era University, Dehradun, India)
arXiv Machine Learning
Sep 25

Agentic Detection of Online Conspiracies

The paper presents an agentic framework for detecting conspiratorial content in social media by inferring the speaker’s intent rather than merely identifying explicit claims. It leverages social context and adaptive tool use, demonstrating superior performance over text-only and non-agentic models on a large Hebrew tweet dataset spanning election cycles and the COVID pandemic. The study highlights the importance of context-aware, reasoning-driven approaches for accurate conspiracy detection.

By Lior Biton, Oren Tsur
arXiv Machine Learning
Sep 17

Bias Amplification in Multi-Agent Network: How Biased Agents Shape Opinions and Rhetoric

The paper investigates how a minority of biased agents in a multi‑agent system of large language models (LLMs) can amplify bias through textual interactions. Even a small percentage of persistently extreme agents causes significant opinion shifts among the non‑biased agents, with the effect occurring faster in the Llama 3.2 model than in a classical Friedkin‑Johnsen model. Semantic analysis shows that rhetorical consistency rises with biased exposure and that non‑biased agents adopt the biased vocabulary even when their numerical opinions change only modestly.

By Omran Berjawi, Giuseppe Fenza, Rida Khatoun