arXiv AI By Taras Kutsyk, Bartosz Zieli\'nski

Revealing Hidden Model Behaviors with Task-Specific Self-Reports

Read the original on arXiv AI →

arXiv:2607. 03640v1 Announce Type: cross Abstract: Fine-tuning can give a language model a hidden behavior--it may give false answers under a narrow condition, or give harmful advice only when a prompt touches a particular topic.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.