arXiv Machine Learning

Online Change-point Detection for Cooperative Multi-Agent Reinforcement Learning

The paper introduces Patterns of Past Rewards (PPR), a lightweight, algorithm‑agnostic online change‑point detector for cooperative multi‑agent reinforcement learning. PPR smooths agents’ return streams, highlights recent changes, and applies a statistical drift detector to flag significant shifts. Experiments in a custom Speaker‑Listener environment show that PPR balances detection speed and alarm stability, outperforming both a smoothed‑return baseline and a raw‑return detector.

arXiv AI
Jul 10

Self-Adaptive Anomaly Detection with Reinforcement Learning and Human Feedback in Connected Vehicles

arXiv:2607. 08373v1 Announce Type: cross Abstract: Connected vehicles are autonomous cyber-physical systems whose behavior must be continuously monitored during operation to detect deviations from normal operation before they propagate into failures.

By Matthias Wei{\ss}, Athreya Hosahalli Prakash, Maurice Artelt, Falk Dettinger, Nasser Jazdi, Michael Weyrich
arXiv Machine Learning
Jul 21

Value-Aware Prediction for Robust Multi-Agent Coordination Under Communication Loss

arXiv:2607. 17914v1 Announce Type: cross Abstract: Robust multi-agent coordination relies heavily on inter-agent communication, which is frequently disrupted by physical and environmental constraints in real-world deployments.

By Kemal Devrim Kafadar, Eren \"Ozaltun, Mahmud Efnan \c{S}anl{\i}, Feyza Orak, Emirhan Gazi, Kubilay Ka\u{g}an K\"om\"urc\"u, Naz{\i}m Kemal \"Ure
Hugging Face Trending Papers
Jul 14

OOD-RL-Bench: A Benchmark Framework for Out-of-Distribution Detection in Reinforcement Learning

Reliable reinforcement learning (RL) agents must maintain operational integrity amidst sensor malfunctions, dynamic disturbances, and slow environmental shifts. The detection of out-of-distribution conditions is pivotal to determining when an agent's observations, transitions, or trajectory dynamics deviate from the assumptions underpinning its policy training.

Hugging Face Trending Papers
Aug 20

End-to-end Early Classification of Time Series in Non-Stationary Environments

Early Classification of Time Series (ECTS) requires making accurate decisions as early as possible in inherently online and evolving environments. Yet, most existing methods assume stationarity and rely on separable designs, where classification and triggering are optimized independently, an assumption that fundamentally limits their adaptability under drift.

arXiv AI
Jun 30

HiComm: Hierarchical Communication for Multi-agent Reinforcement Learning

arXiv:2606. 29126v1 Announce Type: new Abstract: Cooperative multi-agent reinforcement learning (MARL) often relies on communication to mitigate partial observability, yet most existing protocols treat messages as flat dense vectors detached from the structure of the observations they summarize.

By Runze Zhao, Dongruo Zhou, Sumit Kumar Jha, Nathaniel D. Bastian, Ankit Shah