← Back to all news
Hugging Face Trending Papers August 31, 2026

BLOOM-WILT: Logit Tilting for Behaviour Elicitation in Automated LLM Auditing

Read the original on Hugging Face Trending Papers →

The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.

  • llms
  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
3d ago

BLOOM-WILT: Logit Tilting for Behaviour Elicitation in Automated LLM Auditing

arXiv:2608.31105v1 Announce Type: new Abstract: Users of a deployed language model routinely encounter behaviours that testing almost never surfaces, since deployment puts the model through orders of...

By Adrians Skapars, Edoardo Manino
llmssafety
More like this →
arXiv AI
Jun 30

Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives

arXiv:2605. 00994v2 Announce Type: replace-cross Abstract: Finetuning can significantly modify the behavior of large language models, including introducing harmful or unsafe behaviors.

By Mohammed Abu Baker, Luca Baroni, Dan Wilhelm
llmsagentsfine-tuningbenchmarks
More like this →
arXiv AI
2d ago

Self-Reports Are Not Verification: Environment-Grounded Auditing of LLM Operators in Evolutionary Search

arXiv:2609.00652v1 Announce Type: new Abstract: Language model agents increasingly propose actions, observe external feedback, and explain their own behavior. Their confidence and rationales are conv...

By Enrong Pan, Ryan Zhou, Ting Hu
llmsagents
More like this →
arXiv AI
Jun 19

The Autonomy Tax: Defense Training Breaks LLM Agents

arXiv:2603. 19423v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents increasingly rely on external tools (file operations, API calls, database transactions) to autonomously complete complex multi-step tasks.

By Shawn Li, Yue Zhao
llmsagentsbenchmarkssafety
More like this →
arXiv Machine Learning
Jul 1

White-Box Sensitivity Auditing with Steering Vectors

arXiv:2601. 16398v3 Announce Type: replace-cross Abstract: Algorithmic audits are essential tools for examining systems for properties required by regulators or desired by operators.

By Hannah Cyberey, Yangfeng Ji, David Evans
llmssafety
More like this →
arXiv AI
Aug 12

Auditing Automated Evaluation, Error Propagation, and Runtime Mitigation in Tool-Using Language Agents

arXiv:2604. 16706v2 Announce Type: replace Abstract: Automated evaluation of tool-using large language model (LLM) agents is widely assumed to be reliable, yet this assumption is rarely validated against human annotation.

By Bhaskar Gurram
llmsagentsbenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea