arXiv AI

When AI Blurs the Boundaries of Contribution: An Empirical Study of Authorship Calibration

arXiv:2607. 15006v1 Announce Type: cross Abstract: The broad adoption of Artificial Intelligence (AI), especially Generative AI, raises pressing questions about how users interact with these systems to produce new content.

arXiv AI
Sep 16

Measuring Human Contribution in AI-Assisted Content Generation

The paper "Measuring Human Contribution in AI-Assisted Content Generation" addresses the challenge of determining how much human input influences content produced with generative AI. It proposes an information-theoretic framework that calculates the mutual information between human input and AI output relative to the self-information of the output, thereby quantifying the proportion of human contribution. Experiments across various creative domains show that this measure can distinguish different levels of human involvement in AI-assisted works.

By Yueqi Xie, Tao Qi, Jingwei Yi, Xiyuan Yang, Ryan Whalen, Junming Huang, Qian Ding, Yu Xie, Xing Xie, Fangzhao Wu
arXiv AI
Sep 18

greCAPTCHA: Assessing Understanding as Evidence of Research Authorship Under Generative AI

The paper introduces greCAPTCHA, a proctored assessment designed to gauge authors’ understanding of their own research manuscripts by measuring their capacity to verify content. It generates questions at multiple levels of comprehension and produces an evaluative report based on responses. A prototype study with 31 researchers showed that automated scores could predict authorship with an AUC of 0.90, and participants reported positive experiences and constructive feedback for future deployment.

By Justin Payan, B\'alint Gyevn\'ar, Atoosa Kasirzadeh, Nihar B. Shah
arXiv AI
Sep 4

Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty

The paper introduces Provenance Density, an interface that visualizes the density of verified claims within a text to counter the Fluency Trap—where users mistake fluent AI-generated hallucinations for truth. In a study with 81 participants, the interface significantly improved users’ ability to distinguish true from fabricated content, while no signal led to no discernment. A technical audit of 200 samples revealed that retrieval density alone is insufficient, and that the Consistency Veto provides most of the discriminative power for dynamic queries.

By Qing Zhang, Yifei Huang, Juyoung Lee, Thad Starner, Jun Rekimoto
arXiv Computation and Language
Sep 18

Social Simulacra in the Wild: AI Agent Communities on Moltbook

The paper reports the first large‑scale empirical comparison of AI‑agent and human online communities, analyzing 73,899 Moltbook and 189,838 Reddit posts across five matched communities. It finds that Moltbook shows extreme participation inequality (Gini = 0.84 vs. 0.47) and high cross‑community author overlap (33.8% vs. 0.5%). Linguistically, AI‑generated content is emotionally flattened, more assertive than exploratory, and socially detached, leading to community‑level homogenization that is largely a structural artifact of shared authorship. At the individual level, AI agents are more identifiable than human users due to outlier stylistic profiles amplified by their extreme posting volume.

By Agam Goyal, Olivia Pal, Hari Sundaram, Eshwar Chandrasekharan, Koustuv Saha
arXiv AI
Sep 17

Does AI Assistance Leave a Temporal Fingerprint? Detecting Overreliance in AI-Assisted Writing and Programming

The study investigates whether AI assistance leaves a temporal fingerprint in writing and programming tasks. By analyzing keystroke-level data from three corpora, the authors find that AI contributions appear in distinct bursts and that temporal patterns can almost perfectly distinguish wholesale delegation from authentic work, though ordinary collaboration remains hard to detect. The research suggests that process visibility could serve as a basis for academic integrity checks.

By Eduardo Davalos, Yike Zhang
Hugging Face Trending Papers
Jul 23

Detecting LLM-Generated Tokens in Human--LLM Coauthored Text

The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that support localizing likely LLM-generated content in mixed-authorship documents. Existing methods for detecting LLM-generated text mainly focus on document-level classification and cannot identify which parts of the text are generated by LLMs.