arXiv AI By Frank E. Bobe III, Gregory D. Vetaw, Darshan W. Bryner, Matthew G. Cook, Jose L. Salas-Vernis

Deep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models

Read the original on arXiv AI →

Deep Noir is a framework that autonomously discovers optimal activation‑steering parameters in transformer models by leveraging Logit Lens convergence and causal head‑level attribution. It demonstrates significant performance gains across models ranging from 1B to 9B parameters, improving spam detection by up to 42 percentage points and SST‑2 sentiment classification by 13.1 percentage points without code changes. The approach also reveals that increased steering magnitude expands a predictable prompt‑injection attack surface, highlighting security implications for agent systems using steered classifiers.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jul 13

DeepBias: Adaptive In-depth Probing of Social Biases in LVLMs

While Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities, they remain highly susceptible to embedded social biases. Existing bias evaluation protocols predominantly rely on static datasets, which provide only a superficial assessment, as their fixed test cases cannot adaptively evolve to measure the true depth and limits of model vulnerabilities.

arXiv Machine Learning
Jun 11

When is Your LLM Steerable?

arXiv:2606. 11599v1 Announce Type: cross Abstract: Activation steering offers a lightweight approach to control language models' behavior at inference time, but whether it succeeds or fails heavily depends on the prompt, concept, model, and steering configuration.

By Chenrui Fan, Yize Cheng, Ming Li, Soheil Feizi, Tianyi Zhou