arXiv AI

The Evolution of Coordination in a Collective Intelligence System: 25 Years of English Wikipedia and the Emergence of Generative AI

arXiv AI
Jul 1

How Human Feedback Shapes AI-generated Community Notes

arXiv:2606. 30905v1 Announce Type: cross Abstract: Community Notes, a bridging-based crowd-sourced fact-checking system, has emerged as a new mechanism for moderating misleading information on social media and has been adopted by major platforms including X, Facebook, Instagram, Threads, and TikTok.

By Soham De, Isaac Slaughter, Jiawei Guo, Qiao-Yun Cheng, Jiayuan Yan, Sruti Banerjee, Martin Saveski
arXiv Computation and Language
Sep 18

Social Simulacra in the Wild: AI Agent Communities on Moltbook

The paper reports the first large‑scale empirical comparison of AI‑agent and human online communities, analyzing 73,899 Moltbook and 189,838 Reddit posts across five matched communities. It finds that Moltbook shows extreme participation inequality (Gini = 0.84 vs. 0.47) and high cross‑community author overlap (33.8% vs. 0.5%). Linguistically, AI‑generated content is emotionally flattened, more assertive than exploratory, and socially detached, leading to community‑level homogenization that is largely a structural artifact of shared authorship. At the individual level, AI agents are more identifiable than human users due to outlier stylistic profiles amplified by their extreme posting volume.

By Agam Goyal, Olivia Pal, Hari Sundaram, Eshwar Chandrasekharan, Koustuv Saha
arXiv AI
4d ago

The Uneven Decline of Collective Knowledge Production: Evidence from Stack Overflow After Generative AI

The study examines how the release of ChatGPT-3.5 affected collective knowledge on Stack Overflow, analyzing over two million questions from 2020 to 2025. It finds that easy questions declined sharply while difficult ones increased, with rising code complexity. Data-rich topics lost share of questions, whereas data-scarce ones gained, and these patterns held across multiple programming languages.

By Myokyung Han, Taegyoon Kim, Jinhyuk Yun, Lanu Kim
arXiv Computation and Language
Sep 23

Knowledge Pull Requests for Continual Document Authoring

The paper introduces Knowledge Pull Requests (KPRs), a framework that enables continual document authoring by making each change interpretable. KPRs extract claims from new knowledge sources, filter and route them to appropriate sections, and flag conflicts with existing content, producing a ChangeLog that separates knowledge changes from textual edits. Experiments on revising Wikipedia and updating query-driven reports show that KPRs integrate more information, better preserve existing content, and improve question answering performance compared to rewriting from scratch or using frontier models with search.

By Alexander Martin, Benjamin Van Durme
arXiv AI
Aug 28

How LLMs Distort Our Written Language

Large language models (LLMs) are widely used to assist writing, but this study shows they alter both tone and meaning of human text. A user study found that heavy LLM use increased neutral essays by nearly 70% and made writers feel less creative and less in their voice. Even when prompted to make only grammar edits, LLMs changed the semantic content of essays and produced AI-generated scientific reviews that were less focused on clarity and significance and scored higher on average.

By Marwa Abdulhai, Isadora White, Yanming Wan, Ibrahim Qureshi, Joel Z. Leibo, Max Kleiman-Weiner, Natasha Jaques
arXiv Machine Learning
Jun 5

Operation-Guided Progressive Human-to-AI Text Transformation Benchmark for Multi-Granularity AI-Text Detection

arXiv:2606. 06481v1 Announce Type: cross Abstract: As AI writing assistants become increasingly integrated into real-world drafting and revision workflows, many documents are no longer purely human-written or AI-generated, but instead result from progressive human-AI co-editing.

By Sondos Mahmoud Bsharat, Jiacheng Liu, Xiaohan Zhao, Tianjun Yao, Xinyi Shang, Yi Tang, Jiacheng Cui, Ahmed Elhagry, Salwa K. Al Khatib, Hao Li, Salman Khan, Zhiqiang Shen
arXiv AI
Aug 26

Evidence Blindness in Direct Corpus Interaction: Persistent Navigation with AtlasNav

The paper introduces AtlasNav, a persistent multi‑view corpus‑navigation framework that organizes a corpus into a Corpus Atlas, enabling large‑language‑model agents to navigate efficiently under finite interaction budgets. AtlasNav reduces online inference cost by 30.21% and achieves 92.05% strict accuracy on BrowseComp‑Plus, while earlier and more rapidly realizing required evidence compared to dynamic‑workspace methods. The approach also transfers well to other corpora such as PhantomWiki and heterogeneous enterprise knowledge bases, demonstrating that effective agentic search relies on both accessible evidence and a reusable corpus representation.

By Hongyu Guo, Zhiyu Zheng, Zhao Cao
arXiv AI
Sep 25

DocuTeam: Mixed-Initiative Multi-Agent Discussions around Evolving Documents

DocuTeam is a mixed‑initiative multi‑agent discussion system that allows both users and agents to start and steer conversations around evolving documents. Agents monitor changes to the document and proactively initiate or redirect discussions, while users can shape the dialogue or adopt agent suggestions. In a within‑subjects study with 20 participants, DocuTeam produced outcomes that were rated as more novel, relevant, and specific compared to a baseline, without increasing cognitive load.

By Heechan Lee, Juhyeon Choi, Tae Soo Kim, Juho Kim, Joseph Seering