Conformal Policy Control
arXiv:2603. 02196v3 Announce Type: replace Abstract: An agent must try new behaviors to explore and improve.
arXiv:2606. 00320v1 Announce Type: new Abstract: We present an online, distribution-free framework for controlling the Conditional Value-at-Risk (CVaR), extending conformal tail risk control to non-stationary and adversarial environments.
arXiv:2603. 02196v3 Announce Type: replace Abstract: An agent must try new behaviors to explore and improve.
arXiv:2609.38938v1 Announce Type: new Abstract: Reinforcement learning with human feedback (RLHF) learns from human comparisons, which can be corrupted or deliberately manipulated. This paper studies...
arXiv:2505.19893v2 Announce Type: replace Abstract: Large language model pretraining is compute-intensive, yet many tokens contribute marginally to learning, resulting in inefficiency. We introduce E...
arXiv:2605. 14953v2 Announce Type: replace Abstract: We address the problem of conformal selection, where an agent must select a minimal subset of options to ensure that at least one ``success'' is identified with a pre-specified target probability $\phi$.
arXiv:2510. 07750v3 Announce Type: replace-cross Abstract: Robust optimization safeguards decisions against uncertainty by optimizing against worst-case scenarios, yet their effectiveness hinges on a prespecified robustness level that is often chosen ad hoc, leading to either insufficient protection or overly conservative and costly solutions.
The paper introduces a penalized distributionally robust optimization framework that allows an adversary to choose any distribution while incurring a Wasserstein penalty for deviating from the empirical distribution. It shows that the adversary’s problem can be reformulated as optimizing transport maps that push empirical samples to adversarial ones, proving that optimal maps are cyclically monotone. The authors argue that standard per-sample adversarial training violates this property and propose two remedies—multi-start particle ascent and input-convex neural network parameterization—to enforce cyclical monotonicity, demonstrating improved robustness and generalization in experiments on regression, image classification, and control tasks.
arXiv:2607. 24983v1 Announce Type: cross Abstract: Generative models are increasingly adopted in distributionally robust optimization (DRO), but existing approaches trade off model compatibility and adversarial structure: methods that accept arbitrary samplers do not restrict worst-case laws to a generator family, while generator-parameterized adversaries rely on model-specific access such as likelihoods, scores, or training data.
arXiv:2510. 02695v3 Announce Type: replace-cross Abstract: In safety-critical domains where online data collection is infeasible, offline reinforcement learning (RL) is attractive only if policies achieve high returns without catastrophic lower-tail risk.
arXiv:2606. 05551v1 Announce Type: cross Abstract: Reliable decision making pipelines powered by machine learning models require uncertainty quantification (UQ) methods that come with explicit safety guarantees.
arXiv:2609.08064v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in settings where rare but severe harmful generations can have significant consequences. Existin...
arXiv:2602. 01903v2 Announce Type: replace Abstract: This work studies online episodic tabular Markov decision processes (MDPs) with known transitions and develops best-of-both-worlds algorithms that achieve refined data-dependent regret bounds in the adversarial regime and variance-dependent regret bounds in the stochastic regime.
arXiv:2410. 07719v4 Announce Type: replace Abstract: Despite being widely adopted as a canonical framework for learning robust models, adversarial training suffers from robust overfitting.