arXiv Machine Learning By Rachel Freedman

Adaptive Pluralistic Alignment: A pipeline for dynamic artificial democracy

Read the original on arXiv Machine Learning →

arXiv:2605. 01642v2 Announce Type: replace Abstract: Prevailing alignment methods target a fixed set of preferences and therefore risk forcing value lock-in as societal norms evolve over time.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv AI
1d ago

Inference-Time Mitigation of Adversarial Political Bias in Large Language Models

arXiv:2608. 14629v1 Announce Type: cross Abstract: As Large Language Models (LLMs) become the mainstay for information retrieval and summarization tasks, ensuring that they are always non-partisan and invulnerable to political bias is a critical step towards safer and more trustworthy Artificial Intelligence (AI).

By Tejaswi V. Panchagnula, Bruce Coburn, Bryce J. Dietrich, Robert X. Browning, Edward J. Delp, Fengqing Zhu