Hugging Face Trending Papers

SenFlow: Inter-Sentence Flow Modeling for AI-Generated Text Detection in Hybrid Documents

Read the original on Hugging Face Trending Papers →

Sentence-level AI-generated text detection (S-AGTD) for hybrid documents, where humans and LLMs co-author one text, faces two gaps: existing methods classify each sentence in isolation, discarding inter-sentence dependencies, and existing benchmarks omit the newest generation of generators. We construct MOSAIC, a benchmark of 16,000 hybrid documents over PubMed and XSum, generated by DeepSeek-V3.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Machine Learning
Jun 5

Operation-Guided Progressive Human-to-AI Text Transformation Benchmark for Multi-Granularity AI-Text Detection

arXiv:2606. 06481v1 Announce Type: cross Abstract: As AI writing assistants become increasingly integrated into real-world drafting and revision workflows, many documents are no longer purely human-written or AI-generated, but instead result from progressive human-AI co-editing.

By Sondos Mahmoud Bsharat, Jiacheng Liu, Xiaohan Zhao, Tianjun Yao, Xinyi Shang, Yi Tang, Jiacheng Cui, Ahmed Elhagry, Salwa K. Al Khatib, Hao Li, Salman Khan, Zhiqiang Shen
Hugging Face Trending Papers
Aug 27

Relational Over-Regularization: Graph-Based AI-Generated Text Detection via Sentence Transition Deviation

The paper introduces Relational Over-Regularization (ROR), a new way to detect AI‑generated text by examining sentence‑pair transition variance rather than token‑level cues. It proposes the Cross‑Source Stylometric Fingerprint Graph (CSFG), a graph‑based model that encodes positional, sequential, semantic, and transition deviation signals as learnable edge features, achieving 97.14% accuracy on binary detection and outperforming existing graph baselines by 11.14 percentage points. The approach shows strong generalization to unseen large language models in the inflated‑variance regime while maintaining a low false‑positive rate of 1.57%.

arXiv Computation and Language
4d ago

Selecting What Matters: Semantic Compression-Guided Selective Pooling for Long-Context Embeddings

The paper introduces SCSP, a training‑free framework that improves long‑context embeddings by selectively pooling informative tokens. SCSP partitions documents into sentence‑aware chunks, adds a semantic compression prompt to each chunk, and uses prompt‑isolated attention masks to estimate token importance. The selected tokens’ intermediate‑layer representations are aggregated to form the final embedding, yielding consistent performance gains across zero‑shot and fine‑tuned models on long‑context benchmarks.

By Zifeng Cheng, Jie Zheng, Zhiwei Jiang, Shuwen Wang, Fei Shen, Shiping Ge, Qing Gu
arXiv AI
Jun 16

StyleShield: Exposing the Fragility of AIGC Detectors through Continuous Controllable Style Transfer

arXiv:2605. 00924v2 Announce Type: replace-cross Abstract: AI-generated content (AIGC) detectors are increasingly deployed in high-stakes settings such as academic integrity screening, yet their reliability rests on a fundamental paradox: as language models are trained on human-written corpora, the statistical boundary between AI and human writing will inevitably dissolve as models improve.

By Guantian Zheng
arXiv AI
Aug 28

Relational Over-Regularization: Graph-Based AI-Generated Text Detection via Sentence Transition Deviation

The paper introduces Relational Over‑Regularization (ROR), a structural signal at the sentence‑pair level that captures inflated inter‑sentence transition variance in AI‑generated text. It proposes the Cross‑Source Stylometric Fingerprint Graph (CSFG), a graph‑based framework that encodes positional, sequential, semantic, and transition deviation signals as learnable GNN edge features, achieving 97.14% accuracy in binary detection and outperforming existing graph‑based baselines. The method demonstrates robust generalization to unseen large language models in the inflated‑variance regime while maintaining a low false‑positive rate.

By Hyeonchu Park, Bugeun Kim