arXiv:2603. 10938v2 Announce Type: replace-cross Abstract: Safe Reinforcement Learning from Human Feedback (RLHF) typically enforces safety through expected cost constraints, but the expectation captures only a single statistic of the cost distribution and fails to account for distributional uncertainty, particularly under heavy tails or rare catastrophic events.
By Yaswanth Chittepu, Ativ Joshi, Rajarshi Bhattacharjee, Scott Niekum
arXiv:2606. 15531v1 Announce Type: new Abstract: Fine-tuning aligned language models on benign tasks (e.
By Bohdan Turbal, Blossom Metevier, Max Springer, Aleksandra Korolova
arXiv:2608. 08383v1 Announce Type: cross Abstract: Steering vectors are a lightweight tool for controlling LLM behavior.
By Yuxiao Li, Gjergji Kasneci
arXiv:2608. 05045v1 Announce Type: cross Abstract: Released aligned large language models remain vulnerable to malicious downstream finetuning.
By Yuxuan Huang, Xingyu Zeng, Tianhang Zheng, Chaochao Lu
arXiv:2606. 00320v1 Announce Type: new Abstract: We present an online, distribution-free framework for controlling the Conditional Value-at-Risk (CVaR), extending conformal tail risk control to non-stationary and adversarial environments.
By Catherine Chen, Jingyan Shen, Zhun Deng, Lihua Lei
arXiv:2607. 23388v1 Announce Type: cross Abstract: As constrained learning becomes increasingly common, models are trained under explicit feasibility requirements to enforce fairness, safety, robustness, regulariza- tion, and physics or logic constraints.
By Xin Wang (Jeff), R. Tyrrell Rockafellar (Jeff), Xuegang (Jeff), Ban