Pretraining Data Can Be Poisoned through Computational Propaganda
arXiv:2607. 15267v1 Announce Type: new Abstract: Poisoning pretraining data can introduce harmful behaviors to LMs that are difficult to detect and mitigate.
OpenAI researchers collaborated with Georgetown University’s Center for Security and Emerging Technology and the Stanford Internet Observatory to investigate how large language models might be misused for disinformation purposes. The collaboration included an October 2021 workshop bringing together 30 disinformation researchers, machine learning experts, and policy analysts, and culminated in a co-authored report building on more than a year of research.
arXiv:2607. 15267v1 Announce Type: new Abstract: Poisoning pretraining data can introduce harmful behaviors to LMs that are difficult to detect and mitigate.
The paper presents a theory-informed computational framework that converts cross-disciplinary theories of fake news into measurable features for automated detection and explanation. By reviewing theories from social sciences, psychology, economics, and more, the authors establish a broad theoretical foundation for computational modeling. Experiments on benchmark datasets demonstrate that theory-derived features are predictive, provide interpretable diagnostic signals, and that multi-feature models generally outperform individual features, though gains are modest.
arXiv:2609.20838v1 Announce Type: new Abstract: In this study, we examine how modern LLMs generate and detect fake news under controlled settings across four manipulation scenarios. These are open-en...
arXiv:2607. 12336v1 Announce Type: cross Abstract: Artificial Intelligence (AI) technologies, while serving as a foundational enabler for modern social media and digital health services, exert a bivalent effect by simultaneously acting as a combatant against and a spread vector for misinformation.
arXiv:2508.02312v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs), now a foundation in advancing natural language processing, power applications such as text generation, machine...
arXiv:2608. 09510v1 Announce Type: cross Abstract: Detecting machine-generated disinformation on social media is increasingly difficult as large language models (LLMs) make it easier to generate and rewrite misleading content at scale.
arXiv:2608. 15746v1 Announce Type: new Abstract: We present a forensic analysis of the generation pipeline behind a recent AI-driven influence campaign.
arXiv:2607. 14791v1 Announce Type: new Abstract: Transcoders have recently emerged as a promising approach for mechanistic interpretability (MI), enabling circuit-level analysis of model behaviour.
arXiv:2606. 08381v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly released and deployed through opaque development and deployment pipelines, enabling model providers to inject intentional, provider-specific policies without officially announcing them.
arXiv:2607. 10402v1 Announce Type: cross Abstract: Large language models (LLMs) have transformed misinformation from a primarily content-centric problem into a broader ecosystem-level security challenge.
arXiv:2608.21775v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in real-world applications, yet they remain vulnerable to generating harmful content. From adver...
The paper introduces GPTBIAS, a framework that uses powerful large language models like GPT‑4 to evaluate bias in other LLMs. It employs specially crafted prompts called Bias Attack Instructions to probe for bias and outputs a bias score along with detailed information such as bias types, affected demographics, keywords, reasons, and improvement suggestions. Extensive experiments demonstrate the framework’s effectiveness and usability.