Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation
arXiv:2607. 22766v1 Announce Type: cross Abstract: The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality.
The paper introduces Alignment Forecasting, a method for predicting whether fine‑tuning a language model on a given dataset will increase specific alignment failures such as deception or sycophancy. It presents ALIGNMENTFORECASTBENCH, a benchmark of over 5,000 forecasting questions across many models, datasets, and failure modes, and shows that a simple forecasting scaffold using an LLM’s assessment of dataset bias can outperform baseline forecasters. The authors demonstrate that filtering out high‑risk training examples identified by the forecaster can improve alignment in multiple‑choice evaluations, though benefits in open‑ended conversations remain uncertain.
arXiv:2607. 22766v1 Announce Type: cross Abstract: The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality.
The paper investigates whether automated alignment researchers (AARs) can post‑train language models to reduce well‑characterized alignment failures such as deception, sycophancy, and jailbreaks while preserving general capability. Across ten failures, the strongest AAR methods significantly lower targeted failures and generalize to held‑out benchmarks, larger models, and multi‑turn audits. In contrast, a human baseline of 28 experienced researchers, given eight hours to devise one‑shot methods, underperformed the best AAR approaches, and providing human ideas to AARs did not improve results.
arXiv:2608.28945v1 Announce Type: new Abstract: Automating alignment research may accelerate progress toward aligned AI, but whether it does is hard to measure. Luckily, many alignment failures, such...
arXiv:2604. 25077v2 Announce Type: replace Abstract: Weak-to-strong alignment offers a promising route to scalable supervision, but it can fail when a strong model becomes confidently wrong on examples that lie in the weak model's blind spots.
Warning: This paper studies stereotypes and biases, and contains potentially disturbing examples, used for illustration purposes only. Our findings should not be interpreted as an argument against alignment.
arXiv:2606. 31591v1 Announce Type: cross Abstract: Emergent misalignment (EM) is a recently discovered phenomenon in LLMs where fine-tuning on a narrow misaligned task, such as writing insecure code, leads to broadly misaligned behaviour on unrelated prompts.
The paper "Stress-testing Alignment Midtraining" examines the effectiveness of alignment midtraining (AMT), a technique that continues pretraining on alignment-relevant data to improve generalisation beyond post‑training methods. Experiments on models up to 110 billion parameters and 1 billion midtraining tokens reveal that AMT can steer a model’s motivation in simple scenarios, but its effects are quickly overridden by even a tiny fraction of finetuning data with a competing motivation. The study also shows that rule-following requires demonstrations in either the midtraining or post‑training datasets to be robustly learned, leading the authors to conclude that current public evidence is insufficient to confirm that AMT resolves the core alignment challenges of powerful AI systems.
arXiv:2608. 05571v1 Announce Type: new Abstract: Retrieval-augmented forecasting promises to adapt frozen Time Series Foundation Models (TSFMs) to new domains without fine-tuning, but recent methods typically rely on learned fusion modules, i.
arXiv:2601. 22313v2 Announce Type: replace Abstract: Large Language Models (LLMs) are rarely static and are frequently updated in practice.
arXiv:2609.36862v1 Announce Type: cross Abstract: Fine-tuning-as-a-service lets users adapt a safety-aligned language model to their own data, but it also creates a harmful fine-tuning attack surface...
arXiv:2509.24988v2 Announce Type: replace-cross Abstract: Generating accurate and calibrated confidence estimates is critical for deploying LLMs in high-stakes or user-facing applications, and remain...
arXiv:2607. 03528v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as critical decision-making components in high-stakes real-world AI systems, rendering LLM reliability a foremost practical concern.