arXiv:2607. 18263v1 Announce Type: new Abstract: AI-generated non-consensual intimate imagery (AIG-NCII) is not adequately addressed in AI/ML literature regarding AI-generated media, commonly referred to as "deepfakes".
By Li Qiwei, Wells Lucas Santo, Sarita Schoenebeck, Eric Gilbert
The study audits Bluesky’s Moderation Service (BMS) using its 10.6 million public moderation labels from 2025. It finds that BMS operates as a human‑AI collaboration: sexual and graphic content is flagged automatically in seconds, while more nuanced or high‑stakes content requires human review that can take hours or days. The system shows high precision (0.837) but low recall (0.222), with annotators detecting 4.5 times more harmful content than the system, and clustering reveals harms ranging from hostility toward protected groups to the spread of explicit material.
By Pushpdeep Singh, Sayeh Jarollahi, Ayan Majumdar, Vabuk Pahari, Abhijnan Chakraborty, Krishna P. Gummadi, Ingmar Weber, Abhisek Dash
The paper documents a growing trend of removing safety guardrails from open-weight AI models, profiling the ecosystem that produces, redistributes, and applies these models. Between January 2024 and March 2026, 3,471 original uncensored models were identified on HuggingFace, each re‑packaged an average of 2.4 times, with three actors responsible for 52 % of all 8,164 compressed redistributions. After quantization and mirroring across platforms such as Ollama, these models persist even when upstream versions are removed, and 25 % of 1,643 GitHub applications that integrate uncensored large language models were classified as explicitly malicious.
By 10a Labs, :, Juliette Garcia, Hailey May, Bobby McKenzie, David Pham, Matthew Swain, Joshua Valdez, Corie Wieland, Zachary Yahn
arXiv:2605. 08093v2 Announce Type: replace-cross Abstract: The use of chatbots for various forms of companionship is growing rapidly, raising a myriad of questions about simulated relationships, emotional dependence, and psychological harm.
By Maribeth Rauh, Dick A. H. Blankvoort, Matias Duran, Caoilfhionn N\'i Dheor\'ain, Harshvardhan J. Pandit, Syrine Enneifer, Siddharth D. Jaiswal, Anthony Ventresque, Abeba Birhane
arXiv:2606. 22748v2 Announce Type: replace-cross Abstract: Some professional authors are beginning to use AI tools to help produce their fiction writing.
By Neel Gupta, Maria Antoniak, Melanie Walsh
The paper reports the first large‑scale empirical comparison of AI‑agent and human online communities, analyzing 73,899 Moltbook and 189,838 Reddit posts across five matched communities. It finds that Moltbook shows extreme participation inequality (Gini = 0.84 vs. 0.47) and high cross‑community author overlap (33.8% vs. 0.5%). Linguistically, AI‑generated content is emotionally flattened, more assertive than exploratory, and socially detached, leading to community‑level homogenization that is largely a structural artifact of shared authorship. At the individual level, AI agents are more identifiable than human users due to outlier stylistic profiles amplified by their extreme posting volume.
By Agam Goyal, Olivia Pal, Hari Sundaram, Eshwar Chandrasekharan, Koustuv Saha