arXiv Machine Learning By Fatemeh Saberi Khomami, Julita Vassileva

Online Change-point Detection for Cooperative Multi-Agent Reinforcement Learning

Read the original on arXiv Machine Learning →

The paper introduces Patterns of Past Rewards (PPR), a lightweight, algorithm‑agnostic online change‑point detector for cooperative multi‑agent reinforcement learning. PPR smooths agents’ return streams, highlights recent changes, and applies a statistical drift detector to flag significant shifts. Experiments in a custom Speaker‑Listener environment show that PPR balances detection speed and alarm stability, outperforming both a smoothed‑return baseline and a raw‑return detector.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 10

Self-Adaptive Anomaly Detection with Reinforcement Learning and Human Feedback in Connected Vehicles

arXiv:2607. 08373v1 Announce Type: cross Abstract: Connected vehicles are autonomous cyber-physical systems whose behavior must be continuously monitored during operation to detect deviations from normal operation before they propagate into failures.

By Matthias Wei{\ss}, Athreya Hosahalli Prakash, Maurice Artelt, Falk Dettinger, Nasser Jazdi, Michael Weyrich
arXiv Machine Learning
Jul 21

Value-Aware Prediction for Robust Multi-Agent Coordination Under Communication Loss

arXiv:2607. 17914v1 Announce Type: cross Abstract: Robust multi-agent coordination relies heavily on inter-agent communication, which is frequently disrupted by physical and environmental constraints in real-world deployments.

By Kemal Devrim Kafadar, Eren \"Ozaltun, Mahmud Efnan \c{S}anl{\i}, Feyza Orak, Emirhan Gazi, Kubilay Ka\u{g}an K\"om\"urc\"u, Naz{\i}m Kemal \"Ure
Hugging Face Trending Papers
Jul 14

OOD-RL-Bench: A Benchmark Framework for Out-of-Distribution Detection in Reinforcement Learning

Reliable reinforcement learning (RL) agents must maintain operational integrity amidst sensor malfunctions, dynamic disturbances, and slow environmental shifts. The detection of out-of-distribution conditions is pivotal to determining when an agent's observations, transitions, or trajectory dynamics deviate from the assumptions underpinning its policy training.