arXiv:2606. 16219v1 Announce Type: cross Abstract: Digital twin modeling, including control and data assimilation under model uncertainty, often faces an open-ended fidelity problem: adding variables, data streams, and time scales can indefinitely increase model complexity, ultimately producing systems that are difficult to maintain, validate, interpret, and use for stress or safety testing.
By Zongren Zou, Th\'eo Bourdais, Ricardo Baptista, Houman Owhadi
arXiv:2607. 01741v1 Announce Type: cross Abstract: Reinforcement Learning (RL) is a sequential decision-making framework in which an agent learns optimal policies through interaction with an environment by maximizing cumulative rewards.
By Stefano Masini, Cecilia Viscardi, Michela Baccini
arXiv:2607. 18164v1 Announce Type: cross Abstract: Digital Twins rely on surrogate models to mirror physical systems in real time, yet these models can degrade as operating conditions evolve, a phenomenon known as concept drift.
By Yi-Ping Chen, Ying-Kuan Tsai, Vispi Karkaria, Seul Lee, Daniel Apley, Wei Chen
arXiv:2607. 08793v1 Announce Type: cross Abstract: Sepsis is a leading cause of mortality, yet optimal treatment policies remain contested.
By Joshua Pickard, Wei Qi, Na Li, Ann Woolley, Lisa Cosimi, Roy Kishony, Deborah Hung
The paper introduces a closed‑loop cyber‑physical system for autonomous model lifecycle management in automotive manufacturing, deployed since 2023. It manages paired physics and reinforcement‑learning models, selecting the best candidate through competitive retraining cycles and a Conductor orchestrator that handles plant‑wide inventories and fallback controls. The system incorporates an operator‑trust gate that rejects 23% of policies that deviate from established practice, achieving 28‑45% process stability improvements with no safety incidents.
By Zhengyang (Cissy), Gu, Thomas Cook, Fredaljohn Rohrbaugh, Joseph E. Hernandez, Chris Couch
The paper introduces an end‑to‑end model‑based reinforcement learning algorithm that synthesises policies satisfying Linear Temporal Logic (LTL) specifications in unknown environments. It synchronises a Limit‑Deterministic Büchi Automaton (LDBA) with a Bayes‑Adaptive Markov Decision Process (BAMDP) and proposes a novel Bayes‑Adaptive Monte‑Carlo Planning (BAMCP) method for approximate Bayes‑optimal strategy synthesis. Experiments on finite and infinite‑horizon tasks show improved property satisfaction and sample efficiency compared to model‑free baselines, and ablation studies confirm the advantage of the new BAMCP over classical variants, including reduced task violations in cautious RL settings.
By Jonathan Hau, Alessandro Abate