arXiv AI By Lianrong Zuo, Peilan Xu, Yong Liu, Wenjian Luo

Structure-Conditioned Actor-Critic Branches for Quality-Diversity Reinforcement Learning

Read the original on arXiv AI →

arXiv:2606. 08735v1 Announce Type: new Abstract: Quality-diversity reinforcement learning (QD-RL) aims to construct policy repertoires that contain both high-performing and behaviorally diverse policies.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 18

CARE-VI: Conservative Adaptive Reliability Estimation for Value Improvement in Off-Policy Actor-Critic Learning

CARE‑VI introduces a framework for improving value targets in off‑policy actor‑critic learning by combining Conservative Adaptive Ranking and Screening (CARS), Selector‑Evaluator Value Assessment (SEVA), and Dynamic Adaptive Risk‑aware Enhancement (DARE). CARS limits candidate actions to a budgeted prefix and expands it only when uncertainty exceeds a threshold; SEVA orders candidates with selector critics and reviews their values with an evaluator critic, capping the value at the selector reference; DARE adjusts residual corrections based on candidate reliability and signal gaps. Theoretical analysis bounds errors in each component, and empirical tests on SAC, TD3, and TD7 across four MuJoCo tasks show CARE‑VI consistently outperforms baselines in mean return.

By Xiang Zou, Shengzhu Shi, Junqi Gao, Zhichang Guo