arXiv AI

The Uneven Decline of Collective Knowledge Production: Evidence from Stack Overflow After Generative AI

The study examines how the release of ChatGPT-3.5 affected collective knowledge on Stack Overflow, analyzing over two million questions from 2020 to 2025. It finds that easy questions declined sharply while difficult ones increased, with rising code complexity. Data-rich topics lost share of questions, whereas data-scarce ones gained, and these patterns held across multiple programming languages.

arXiv Computation and Language
Sep 4

Beyond Accuracy: Community Perspectives on Machine Translation

The paper examines how four stakeholder groups—AI developers, professional translators, language learners, and language service providers—discuss machine translation on social media. Using a dataset of 79,286 posts from Reddit, Facebook, Bluesky, and Mastodon (2019‑2025), the authors find frequent disagreements and strong conflicts over translation quality, efficiency, and reliability. These conflicts arise because AI communities view the issues as technical, while non‑AI users prioritize quality nuances, time savings, trust, and broader social concerns.

By Yujun Wang, Ehud Reiter, Shimei Pan, Steffen Eger, Wei Zhao
arXiv Computation and Language
Sep 18

Finding Common Ground: Graded Communal Knowledge in Bluesky Starter Packs

The paper investigates how shared community affiliations, measured via Bluesky starter packs, correlate with common ground between users. By analyzing 191,648 user pairs, it finds that lexical similarity—used as a proxy for common ground—increases monotonically with the number of shared starter packs, especially when those packs represent distinct topical communities. The study also shows that this effect is independent of network proximity, indicating that community membership is a distinct, measurable carrier of common ground.

By Sagar Kumar, Lawrence Swaminathan Xavier Prince, Julia Mendelsohn, Brooke Foucault Welles, Nicholas W. Landry
Hugging Face Trending Papers
Jun 8

Beyond Accuracy: Community Perspectives on Machine Translation

Despite remarkable progress in machine translation (MT), non-AI communities have raised growing concerns about MT systems, suggesting a noticeable gap between technical advancement and the needs of real-world users. For instance, while NLP researchers focus on benchmark performance, end users care about ethical concerns, trust, reliability, costs, and more.

arXiv AI
Jun 2

Characterizing Web Search in The Age of Generative AI

arXiv:2510. 11560v2 Announce Type: replace-cross Abstract: The advent of LLMs has given rise to generative search, a new search paradigm in which LLMs retrieve information from the web related to a query and synthesize it into a single, coherent response.

By Elisabeth Kirsten, Jost Grosse Perdekamp, Qinyuan Wu, Mihir Upadhyay, Krishna P. Gummadi, Muhammad Bilal Zafar
arXiv AI
Aug 5

Towards a new paradigm of scientific discovery with socialized artificial intelligence

arXiv:2608. 02775v1 Announce Type: new Abstract: Scientific discovery has advanced through successive transformations in the organization of knowledge.

By Xinjie Yao, Xingxin Xu, Xiyuan Gao, Zhoupeng Guo, Kunlong Yang, Dengyu Zhao, Siqi Zhao, Zhihe Fan, Yichen Dong, Xin Li, Jiekang Feng, Jiahe Wu, Sen Wang, Beiming Yu, Kejia Zhao, Ruipu Zhao, Jiaqi Zhou, Heyang Li, Jianjun Chen, Anbo Dai, Xin Liu, Zhengtao Yu, Qinghua Hu, Pengfei Zhu
arXiv Computation and Language
Sep 18

Social Simulacra in the Wild: AI Agent Communities on Moltbook

The paper reports the first large‑scale empirical comparison of AI‑agent and human online communities, analyzing 73,899 Moltbook and 189,838 Reddit posts across five matched communities. It finds that Moltbook shows extreme participation inequality (Gini = 0.84 vs. 0.47) and high cross‑community author overlap (33.8% vs. 0.5%). Linguistically, AI‑generated content is emotionally flattened, more assertive than exploratory, and socially detached, leading to community‑level homogenization that is largely a structural artifact of shared authorship. At the individual level, AI agents are more identifiable than human users due to outlier stylistic profiles amplified by their extreme posting volume.

By Agam Goyal, Olivia Pal, Hari Sundaram, Eshwar Chandrasekharan, Koustuv Saha