Decision Making Needs Uncertainty Quantification [Lecture Notes]
arXiv:2607. 14407v1 Announce Type: cross Abstract: Many signal processing systems ultimately exist to {act}.
arXiv:2606. 08552v1 Announce Type: new Abstract: I discuss some quantitative representations of Promise Theory for processes involving autonomous agents.
arXiv:2607. 14407v1 Announce Type: cross Abstract: Many signal processing systems ultimately exist to {act}.
arXiv:2606. 07017v1 Announce Type: new Abstract: Foundation model agents are increasingly deployed for real-world decision-making, but suffer from the sim-to-real gap.
As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified. We present a formal safety argument for the Scientist AI (SAI) Predictor, trained to approximate the Bayesian posterior conditioned on a dataset of "epistemically contextualized" natural-language statements.
arXiv:2306. 02704v2 Announce Type: replace-cross Abstract: We introduce \emph{Calibrated Stackelberg Games (CSGs)}, a generalization of the standard Stackelberg Games (SGs) framework.
arXiv:2606. 29657v1 Announce Type: new Abstract: As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified.
arXiv:2606. 18746v1 Announce Type: new Abstract: This paper develops a formal account of what generalist agents must store in memory in order to act near-optimally across multiple environments and goals.
Eureka is a task‑conditioned Meta‑Agent architecture that transforms long‑horizon scientific tasks into dynamic obligation graphs with explicit acceptance semantics. During execution it constructs Macro‑Agents equipped with specialized state, memory, operators, tools, verifiers, and local topology, using receding‑horizon planning, architecture promotion, and minimal‑sufficient compilation. The system demonstrates strong empirical performance, completing all 170 recursive tasks, generating 3,948 certificates without false acceptances, and achieving significant reductions in input size, recomputation, and consistent serialization across 16,000 concurrent executions.
arXiv:2606. 24842v1 Announce Type: new Abstract: In the big-world regime, agents cannot be universally capable and their ability is inevitably specialized across a world model in pieces.
The paper introduces a Bayesian self‑escalation strategy for hierarchical large‑language‑model agents, allowing an agent to detect during its own reasoning that it is unlikely to succeed and hand control over to a stronger model. The authors formalise this as an optimal‑stopping problem over a learned competence posterior, derive a myopic escalation threshold, and prove that the optimal policy is a time‑varying threshold without assumptions on the raw signal. They provide theoretical guarantees—including a 1/√n regret decay with n calibration trajectories—and validate the approach in simulations and a real‑model code‑generation cascade, showing that the escalation frontier outperforms post‑hoc routing at equal cost. whyItMatters":"The study offers a principled, theoretically grounded method for agents to dynamically decide when to seek stronger models, potentially improving efficiency and reliability in hierarchical LLM systems."
arXiv:2511. 22226v2 Announce Type: replace Abstract: The standard theory of model-free reinforcement learning assumes that the environment dynamics are stationary and that agents are decoupled from their environment, such that policies are treated as being separate from the world they inhabit.
arXiv:2601. 19612v3 Announce Type: replace-cross Abstract: Safe exploration is a key requirement for reinforcement learning (RL) agents to learn and adapt online, beyond controlled (e.
arXiv:2606. 05551v1 Announce Type: cross Abstract: Reliable decision making pipelines powered by machine learning models require uncertainty quantification (UQ) methods that come with explicit safety guarantees.