arXiv AI By Bulambo Mwendelwa Gloire, Prasenjit Mitra

Large Language Models Threaten Double-blind Review

Read the original on arXiv AI →

arXiv:2608. 05157v1 Announce Type: cross Abstract: Double blind peer review serves as the scientific community primary defense against status and affiliation bias.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jun 11

Authority, Truth, and Citation Bias: A Large-Scale Multi-Domain Benchmark for Studying Epistemic Susceptibility in Large Language Models

Large language models are increasingly deployed in citation-augmented settings, yet the effect of citation presence on model behavior independent of factual content remains poorly understood. We introduce AuthorityBench, a 220,564-prompt multi-domain benchmark that isolates how citation-based authority signals influence epistemic behavior in LLMs.

arXiv AI
6d ago

Attribution Bias in Large Language Models

The paper introduces AttriBench, a benchmark dataset that balances author fame and demographics to study quote attribution in large language models (LLMs). Using AttriBench, the authors evaluate 11 popular LLMs and find that accurate attribution remains difficult, with significant disparities across race, gender, and intersectional groups. They also identify a new failure mode—suppression—where models omit attribution entirely, which is unevenly distributed across demographics and not reflected by standard accuracy metrics.

By Eliza Berman, Bella Chang, Daniel B. Neill, Emily Black
Hugging Face Trending Papers
Aug 17

Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies

Can a language model recover the true research idea of a published paper when given only that paper's pre-publication bibliography? We introduce Reconstruction, a blind idea-recovery benchmark that withholds the seed paper and all contemporaneous or future literature, and asks models to propose hypotheses that an independent large language model judge matches against the held-out ground-truth idea.

arXiv AI
Aug 20

Self-prompting and cross-model consensus enable reproducible data extraction from scientific literature with large language models

The paper investigates how large language models can extract contextualized data from scientific literature. It presents four workflows: expert‑written prompts, self‑generated prompts, autonomous literature discovery, and dataset creation from guidelines. While models perform well with prompts, they struggle with context, hallucinate references, and still need human oversight for final validation.

By Valentin Romanov, Monique Bax, Steven Niederer