arXiv AI By Samira Hajizadeh

Retroactive Chain-of-Thought (RetroCoT): Forensic Reconstruction Prompts as a Safety Diagnostic Across Model Generations

Read the original on arXiv AI →

arXiv:2607. 04645v1 Announce Type: cross Abstract: Safety alignment in large language models is typically evaluated against direct, imperative harmful requests.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 2

Can LLMs Reliably Self-Report Adversarial Prefills, and How?

The study investigates whether large language models (LLMs) can reliably detect when their own responses have been manipulated by adversarial prefill attacks. Across ten instruction‑tuned LLMs ranging from 3B to 70B parameters and four safety benchmarks, none consistently recognized compromised outputs, with models claiming intent on prefilled responses at an average of 25.3%. The research identifies that introspective signals mainly arise from safety reasoning and refusal, and that training to improve introspection can paradoxically increase attack success, underscoring the fragility of LLM self‑reporting in safety contexts.

By Quang Minh Nguyen, Uzair Ahmed, Taegyoon Kim