arXiv AI By Samira Hajizadeh

Retroactive Chain-of-Thought (RetroCoT): Forensic Reconstruction Prompts as a Safety Diagnostic Across Model Generations

Read the original on arXiv AI →

arXiv:2607. 04645v1 Announce Type: cross Abstract: Safety alignment in large language models is typically evaluated against direct, imperative harmful requests.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.