Decision Making Needs Uncertainty Quantification [Lecture Notes]
arXiv:2607. 14407v1 Announce Type: cross Abstract: Many signal processing systems ultimately exist to {act}.
arXiv:2608. 06420v1 Announce Type: cross Abstract: Perception in biological systems is inherently noisy, requiring organisms to make decisions under uncertainty where misclassification can be costly or fatal.
arXiv:2607. 14407v1 Announce Type: cross Abstract: Many signal processing systems ultimately exist to {act}.
arXiv:2606. 10228v1 Announce Type: cross Abstract: Safe exploration is a prerequisite for deploying reinforcement learning (RL) agents in safety-critical domains.
Intrinsic motivation plays a central role in adaptive and goal-directed behavior by conferring agents reward-independent objectives and biases useful to act in noisy and uncertain environments. Active...
arXiv:2609.39342v1 Announce Type: cross Abstract: Intrinsic motivation plays a central role in adaptive and goal-directed behavior by conferring agents reward-independent objectives and biases useful...
The paper introduces RATTL (Risk-Adversarial Total-Reward Learning), a framework that adjusts an agent’s caution based on epistemic uncertainty by using a Bayesian posterior over dynamics and a Wasserstein ambiguity set whose radius depends on that posterior. As evidence accumulates, the radius shrinks, smoothly transitioning the agent’s behavior from worst-case robustness to risk-neutral reward maximization. The authors prove a Safety Sandwich theorem showing RATTL’s value lies between the uninformed robust value and the full-knowledge optimum, and demonstrate the method on a binary-hazard example where the criterion reduces to Conditional Value-at-Risk.
arXiv:2608. 04232v1 Announce Type: new Abstract: Biological systems must regulate competing needs under limited perceptual bandwidth, where sharpening one estimate costs the capacity to sharpen the others.
The paper introduces Interval POMDP Shielding for agents with imperfect perception, aiming to prevent unsafe actions when sensor readings may be misclassified. By estimating perception uncertainty from finite labeled data, the authors construct confidence intervals and model the system as a finite Interval Partially Observable Markov Decision Process. They propose an algorithm that computes a conservative belief set, enabling a runtime shield that guarantees, with high probability, that any action allowed by the shield meets a specified safety lower bound. Experiments on four case studies demonstrate that this shielding approach outperforms state‑of‑the‑art baselines in safety.
The paper investigates when intrinsic rewards effectively drive exploration in reinforcement learning. It introduces a formal criterion that evaluates policies based on the counterfactual information they acquire, comparing how well their histories can replace experience from alternative policies. Using a simple environment, the authors show that common intrinsic reward objectives—count-based, prediction-error, empowerment, and information-gain—can lead to Pareto-suboptimal exploration under this criterion, and they propose conditions and a new objective that better align with optimal exploration.
arXiv:2606. 06976v1 Announce Type: new Abstract: Large language model (LLM)-based agents often make suboptimal tool-use decisions, including unsupported tool invocation and hallucinated direct responses, which may accumulate errors throughout multi-step interactions.
arXiv:2510. 16462v3 Announce Type: replace Abstract: This work introduces MAYA, a sequential imitation learning model based on multi-armed bandits, designed to reproduce and predict individual bees' decisions in contextualized foraging tasks.
arXiv:2606. 30136v1 Announce Type: new Abstract: Humans facing algorithmic decision systems have been found to ``game'' them by altering their input data (at a cost to them) in order to favorably change the algorithmic outcomes they receive (at a cost to the algorithm).
arXiv:2507.11482v5 Announce Type: replace Abstract: Artificial learning systems are graduating from passive learners to increasingly autonomous agents, lending pragmatic urgency to the question of wh...