Values as Style: Disentangling Values from Semantics with One-Way Mixing for Low-Damage LLM Steering
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2602. 07356v2 Announce Type: replace Abstract: Aligning large language models (LLMs) with human values has become increasingly important as their influence on human behavior and decision-making expands.
arXiv:2609.39929v1 Announce Type: cross Abstract: Long-context failures of RoPE-based language models can arise from RoPE's intrinsic tradeoff between maintaining stable token preferences and disting...
The paper introduces ALOE, a method for knowledge editing that learns semantic addresses from paraphrases and hard negatives, aligns them with autoregressive hidden states, and embeds a gated low‑rank operator within a single MLP layer. This design allows the edited model to run in one forward pass without external retrievers or routers. Experiments on CounterFact, ZSRE, and KnowEdit show high efficacy (0.955–0.999) and locality (0.981–1.000) across 7–8B model families, with analyses indicating effective separation of edits and suppression of out‑of‑scope activation.
Long-context failures of RoPE-based language models can arise from RoPE's intrinsic tradeoff between maintaining stable token preferences and distinguishing nearby positions. Determining which weaknes...
arXiv:2602. 03160v2 Announce Type: replace Abstract: Aligning Large Language Models (LLMs) with the diverse spectrum of human values remains a central challenge: preference-based methods often fail to capture deeper motivational principles.
arXiv:2507. 18043v2 Announce Type: replace-cross Abstract: Inference-time steering methods offer a lightweight alternative to fine-tuning large language models (LLMs) and vision-language models (VLMs) by modifying internal activations at test time without updating model weights.