← Back to all news
Hugging Face Trending Papers September 1, 2026

Trust Your Guide Only When Certain: Uncertainty-Aware Sparse Alignment at Inference Time

Read the original on Hugging Face Trending Papers →

The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.

  • llms
  • benchmarks
  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Computation and Language
Sep 2

Trust Your Guide Only When Certain: Uncertainty-Aware Sparse Alignment at Inference Time

arXiv:2609.00624v1 Announce Type: new Abstract: A prominent paradigm in inference-time alignment employs lightweight supervisors to steer Large Language Models (LLMs). Through empirical analysis, we...

By Zeen Zhu, Zhuo Li, Weiyang Guo, Liye Zhao, Haibing Di, Yequan Wang, Jing Li
llmsbenchmarkssafety
More like this →
arXiv Machine Learning
Jun 4

Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuning

arXiv:2603. 07445v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often require fine-tuning (FT) to perform well on downstream tasks, but FT can induce safety-alignment drift even when the training dataset contains only benign data.

By Guoli Wang, Haonan Shi, Tu Ouyang, An Wang
llmsfine-tuningsafety
More like this →
Hugging Face Trending Papers
Jun 24

PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models

Safety alignment of large language models (LLMs) typically depends on high-quality supervision data, such as safe demonstrations or preference pairs. However, in real-world deployment, emerging safety requirements are often specified as natural-language policies, while corresponding supervision data may be costly, delayed, or unavailable.

llmsefficiencysafety
More like this →
arXiv AI
Jun 11

To Intervene or Not: Guiding Inference-time Alignment with Probabilistic Model Blending

arXiv:2606. 11201v1 Announce Type: cross Abstract: The wide deployment of LLMs has made model alignment necessary to make newly trained models safely and effectively respond to user instructions.

By Jin Gan, Xin Li, Jun Luo
llmssafety
More like this →
arXiv AI
Jun 2

SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment

arXiv:2606. 02530v1 Announce Type: new Abstract: Aligning Large Language Models (LLMs) with human values often degrades their general capabilities, termed the alignment tax.

By Hao Li, Jingkun An, Zijun Song, Pengyu Zhu, Rui Li, Hao Wang, Wendi Feng, Yesheng Liu, Lijun Li, Jin-Ge Yao, Lei Sha
llmsefficiencybenchmarkssafety
More like this →
arXiv AI
Sep 1

Beyond Token-Level Guidance: Inference-Time Alignment of Specialized LLMs via Cross-Family Representation Steering

arXiv:2608.30319v1 Announce Type: cross Abstract: Large language models (LLMs) finetuned for specialized domains represent crucial high-impact applications. Inference-time alignment improves safety d...

By Jin Gan, Xin Li, Jun Luo
llmsfine-tuningbenchmarkssafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea