Hugging Face Trending Papers

Factions Within, Uncertain Across: Within-Document Reader Sub-Groups in Social Highlighting

When many people highlight the same document, is the crowd a single consensus, or is it internally structured into reader sub-groups that mark different things -- and is that structure a stable property of a reader or of the document? Building on prior work showing an individual's within-document highlighting signal is a whisper while individuality lives in selection, we ask the group-level question on a co-readership platform using a margin-preserving curveball null.

Hugging Face Trending Papers
Aug 3

Floor, Ceiling, and the Fusion Gap: How Much of Crowd Reading Attention Can Machines Predict?

A benchmark score means nothing without knowing what a trivial method achieves and what the best possible method could achieve. We construct both bounds for a task with a rare kind of ground truth: predicting which sentences a crowd of readers -- highlighting for their own purposes, unpaid, uninstructed, and blind to each other -- marked in 120 web documents.

Hugging Face Trending Papers
Aug 19

Readable, Faithful, Used: Three Dissociable Properties of Demographic Identity in a Language Model

The study investigates how demographic identity is represented in a language model, using representational similarity analysis against Pew survey data across 169 demographic cells. It finds that standard last‑token read‑outs underestimate the model’s fidelity, while specific attention heads (notably L11 H16) capture demographic structure more accurately, though race‑based types remain weak. Causal interventions reveal that high fidelity does not guarantee causal use, and a 128‑dimensional probe of a single head improves alignment with survey truth but fails to recover per‑question group ordering.

arXiv AI
6d ago

Peer-Voted LLM-Agent Stress Tests Find Feed-Induced Lexical Convergence but No Reliable Matched-Exposure Advantage for Distributed Sources

The study introduces PV‑SST, a peer‑voted social‑platform testbed, and conducts a preregistered matched‑exposure experiment across four topics, four seeds, four model families, and three larger variants, totaling 448 trials. Results show that feeding agents a ranked list of prior peer posts increases lexical similarity in both the core panel and larger variants, but does not produce a reliable advantage in opinion alignment or survival rates. The only robust finding is lexical convergence driven by the peer‑ranked feed, with no consistent coordination benefit across models or topics.

By Rana Muhammad Usman, Dominic Williamson
arXiv AI
Aug 20

Decomposing Wrong-Consensus Agreement in LLM Self-Consistency: A GPT-4.1 Case Study

The paper introduces a pluralistic agreement index, Gamma, to quantify how often wrong runs of large language models (LLMs) agree with the majority consensus. By decomposing Gamma into a mechanical component and a preference‑unexplained residual, the authors show that on GPT‑4.1 the mechanical part explains most of the agreement on multiple‑choice benchmarks but only about half on open‑domain tasks, revealing a residual bias that can cause self‑consistency to backfire on hard questions. The study provides a quantitative framework for understanding when majority voting over LLM samples improves or harms accuracy, without proposing new voting methods.

By Lizhuo Zhang, Mengmeng Tang, Chenfeng Long, Xiaoyong Tang, Xiang Luo
arXiv AI
6d ago

Structure for Reading, Prose for Writing: Asymmetric Structural Conditioning in Multi-Agent Document Authoring

The paper reports on a deployed multi‑agent tender‑response system that uses an open‑weights language model under sovereignty constraints. In a blind comparison, the system’s answers were judged at least as good as human‑written bids in 40 of 55 sections, with only a few gaps attributable to missing knowledge rather than writing quality. The study also demonstrates an asymmetry in conditioning: while structural markup improves reading tasks, converting instruction material from prose to nested XML degrades answer quality, and naming forbidden constructions concentrates defects.

By Cheng Yu, Nikhil Mathew, Zhengjie Wang
Hugging Face Trending Papers
Aug 18

Preference Is Not Intervention: The Structure and Stability Boundaries of Reader-Specific Evidence Utility

The paper investigates whether reader-specific differences in retrieval‑augmented generation (RAG) reflect reusable structure or merely input‑local interactions. By fixing query, evidence, task, scoring, and intervention, the authors find that nine readers disagree on the effect sign in 33% of cases, with reader×query interactions explaining 29.8% of utility variance. They further decompose heterogeneity into evidence activity, ordinal preference, and conditional signed direction, discovering that ordinal reader geometry is stable across multiple settings while signed geometry is task‑bounded, yet stable ordinal similarity does not predict cross‑reader intervention transfer.

arXiv Computation and Language
5d ago

Flesch-Kincaid Readability Depends Only on the Topic Distribution in Long Texts under Topic Models

The paper shows that the Flesch Reading Ease and Flesch‑Kincaid Grade Level scores, which are computed from the same two document statistics, converge almost surely to deterministic functions of a document’s topic distribution when modeled with a topic model that includes explicit sentence boundaries. In the long‑text limit, all variation in these scores is driven solely by topical composition, not by any residual readability signal. Experiments on the Brown and BNC corpora demonstrate that a topic vector inferred from one half of a document can predict the other half’s FKGL with substantial correlation (r = 0.779 and 0.884), though adding this prediction to genre and syllable‑count features yields only marginal gains in explained variance.

By Yo Ehara