arXiv AI By Juergen Dietrich

Can Multi-Agent LLMs Identify Their Peers? Stylometric Fingerprinting in Role-Constrained Political Analysis

Read the original on arXiv AI →

arXiv:2606. 09854v1 Announce Type: cross Abstract: Multi-agent large language model (LLM) pipelines for political statement analysis are vulnerable to peer-preservation bias: models tend to protect peer models from deactivation and show identity-dependent scoring distortions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 10

IDP-Bench: Benchmarking ability of LLMs to protect personal information in interdependent privacy contexts

arXiv:2606. 09908v1 Announce Type: cross Abstract: Large language models (LLMs) are becoming widely deployed as personal AI assistants with access to sensitive user data, making privacy a major challenge for their design and evaluation.

By Ayana Hussain, Soumya Sharma, Golnoosh Farnadi, Nicholas Vincent, H\'eber Hwang Arcolezi, Ulrich A\"ivodji
arXiv Machine Learning
Sep 22

Fairness Beyond Anonymization? Demographic Leakage in German LLM-Generated Resumes

The study audits demographic leakage in German-language resumes generated by large language models. Using ChatGPT, Gemini, and Qwen 3 variants, the authors generate resumes from anonymized profiles, varying only gender- and ethnicity-associated names while keeping qualifications constant. Even after anonymization and gender-neutralization, classifiers can reliably distinguish male- from female-generated resumes, driven by subtle differences in gender-neutral terminology rather than overtly gendered wording; ethnicity-related leakage remains weak.

By Charlotte Leininger, Helena Veit, Matthias A{\ss}enmacher, Andreas Bender
arXiv AI
Sep 7

MABPD: Multi-Agent Bias Probing & Detection via Structured Argument Debate

MABPD (Multi‑Agent Bias Probing & Detection) is a training‑free pipeline that uses three specialized large language model agents to analyze news articles from complementary perspectives and resolve disagreements via a Structured Argument Debate (SAD) protocol. SAD imposes an asymmetric burden of proof—biased claims lacking grounded textual evidence receive zero weight—along with role‑weighted voting and post‑consensus verification, replacing task‑specific supervised decision boundaries. Ablation studies show that the debate module alone accounts for up to a 10.6‑point F1 gain, and on the BABE benchmark MABPD attains 83.4% macro F1, within 0.7 percentage points of the supervised state‑of‑the‑art, while achieving 75.0% zero‑shot accuracy on the SemEval 2019 HyperPartisan corpus.

By Garvit Joshi (Graphic Era University, Dehradun, India), Stavya Dhyani (Graphic Era University, Dehradun, India), Jasmine (Graphic Era University, Dehradun, India), Arun Chauhan (Graphic Era University, Dehradun, India)
arXiv AI
Aug 11

Knowing You Is Everything: LLM Agents Achieve Near-Perfect Profile-Consistent Reaction Prediction in Social Media Simulation

arXiv:2608. 07498v1 Announce Type: cross Abstract: Autonomous AI agents in social media present concrete risks to democratic discourse and platform governance, while also offering tools for pre-deployment recommender system testing.

By Ljubisa Bojic, Ljiljana Matic, Joerg Matthes, Milan Cabarkapa, Bojana Dinic, Jue Wang