arXiv AI

Uncensored Open-weight Models: Redistribution as the Persistence Layer

The paper documents a growing trend of removing safety guardrails from open-weight AI models, profiling the ecosystem that produces, redistributes, and applies these models. Between January 2024 and March 2026, 3,471 original uncensored models were identified on HuggingFace, each re‑packaged an average of 2.4 times, with three actors responsible for 52 % of all 8,164 compressed redistributions. After quantization and mirroring across platforms such as Ollama, these models persist even when upstream versions are removed, and 25 % of 1,643 GitHub applications that integrate uncensored large language models were classified as explicitly malicious.

arXiv Machine Learning
Jul 14

One Token Is Enough: Fingerprinting and Verifying Large Language Models from Single-Token Output Distributions

arXiv:2607. 10252v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly consumed through opaque serving chains - API aggregators, resellers, and inference providers - in which the client has no technical means to confirm that the model answering is the model advertised, and recent audits show that a substantial fraction of commercial endpoints deviate from the vendor's reference weights.

By Tomas Bruckner
arXiv AI
6d ago

Cheap, open agents make LLM pollution harder to mitigate

The paper titled "Cheap, open agents make LLM pollution harder to mitigate" reports that open-weight language models combined with open-source agentic frameworks can produce synthetic survey responses that are competitive with commercial agents and harder to detect. The authors compared nine agent configurations, finding that fully open agents run locally without usage fees and that no single detection check reliably identifies all agents. Open-text responses were the most effective at distinguishing agents from humans, highlighting the need for multilayered detection strategies.

By Raluca Rilla, Anne-Marie Nussberger, Rui Mata, Dirk U. Wulff
arXiv Machine Learning
Jun 2

Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models

arXiv:2606. 00566v1 Announce Type: new Abstract: As language models take on agentic roles that span calling external APIs, reading tool outputs, and acting on instructions embedded in third-party content, their attack surface expands well beyond what users type.

By Mohammed Sameer Syed (University of Arizona), Rozhin Yasaei (University of Arizona)