CARE-VI: Conservative Adaptive Reliability Estimation for Value Improvement in Off-Policy Actor-Critic Learning
Read the original on arXiv Machine Learning →CARE‑VI introduces a framework for improving value targets in off‑policy actor‑critic learning by combining Conservative Adaptive Ranking and Screening (CARS), Selector‑Evaluator Value Assessment (SEVA), and Dynamic Adaptive Risk‑aware Enhancement (DARE). CARS limits candidate actions to a budgeted prefix and expands it only when uncertainty exceeds a threshold; SEVA orders candidates with selector critics and reviews their values with an evaluator critic, capping the value at the selector reference; DARE adjusts residual corrections based on candidate reliability and signal gaps. Theoretical analysis bounds errors in each component, and empirical tests on SAC, TD3, and TD7 across four MuJoCo tasks show CARE‑VI consistently outperforms baselines in mean return.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.