Explainable Reinforcement Learning for Adaptive Traffic Signal Control
arXiv:2607. 03703v1 Announce Type: new Abstract: Reinforcement Learning (RL) has emerged as a powerful paradigm for adaptive traffic signal control.
SIGMA is a reinforcement‑learning framework for traffic signal control that incorporates a large language model to adaptively tune multiple objectives based on natural‑language emergency commands. It uses rotational data augmentation to learn orientation‑invariant policies and an offline‑to‑online training pipeline to ensure stable deployment. Experiments in SUMO on four Kolkata intersections show that SIGMA reduces waiting times, queue lengths, and improves throughput compared to fixed‑time, actuated, and DQN baselines, with ablation studies confirming robustness to component failures and geometric rotations.
arXiv:2607. 03703v1 Announce Type: new Abstract: Reinforcement Learning (RL) has emerged as a powerful paradigm for adaptive traffic signal control.
The paper presents a modeling and simulation framework to study reinforcement‑learning control of connected and automated vehicle (CAV) platoon joining maneuvers in mixed traffic. It evaluates Deep Q‑Network, Double Deep Q‑Network, and Proximal Policy Optimization algorithms, finding that PPO achieves a 98 % joining success rate with less than 1 % collisions by incorporating risk penalties, though it requires more decision steps. An external safety controller can prevent collisions but may reduce joining efficiency, highlighting a trade‑off between safety, effectiveness, and decision speed.
arXiv:2601. 18783v2 Announce Type: replace-cross Abstract: Balancing safety, efficiency, and operational costs in highway driving poses a challenging decision-making problem for heavy-duty vehicles.
The paper presents a modeling and simulation framework to study reinforcement learning (RL) control of connected and automated vehicle (CAV) platoon joining maneuvers in mixed traffic. Using SUMO and agent-based modeling, it evaluates Deep Q-Network (DQN), Double DQN (DDQN), and Proximal Policy Optimization (PPO) algorithms, finding that PPO achieves a 98 % joining success rate with less than 1 % collision rate by incorporating risk penalties. The study also shows a trade‑off between safety, joining effectiveness, and decision efficiency, and demonstrates that an external safety controller can prevent collisions but may reduce joining efficiency.
arXiv:2609.36934v1 Announce Type: new Abstract: Traffic signal control (TSC) is essential for improving urban mobility and reducing congestion. Although roadside cameras are widely deployed at signal...
arXiv:2607. 18286v1 Announce Type: cross Abstract: Transit signal priority (TSP) requires balancing competing objectives: reducing bus delay while limiting adverse impacts on non-bus traffic and avoiding extreme waits for a subset of vehicles.
arXiv:2412.02520v4 Announce Type: replace-cross Abstract: Connected automated vehicles (CAVs) equipped with adaptive cruise control (ACC) create new opportunities for highway congestion mitigation. T...
arXiv:2604. 17456v2 Announce Type: replace Abstract: Large language model (LLM) agents have shown strong capabilities in long-horizon reasoning, tool use, and decision-making in digital environments, yet extending them to physically grounded systems remains challenging.
arXiv:2606. 26527v1 Announce Type: new Abstract: Transfer learning improves policy learning efficiency by reusing knowledge from source tasks, providing a feasible paradigm for safe and efficient autonomous highway lane changing decision-making.
arXiv:2603. 06607v2 Announce Type: replace-cross Abstract: Radio resource allocation (RRA) is a critical function in cellular vehicle-to-everything (C-V2X) networks, where vehicles must share limited wireless resources to support safety-critical communications.
The paper introduces WM‑RMoE, a World Model‑based Risk‑aware Mixture‑of‑Experts framework for autonomous overtaking. It uses a learned latent dynamics model to perform multi‑step rollouts, evaluating cumulative risk at the trajectory level, and employs a hierarchical gating mechanism to coordinate long‑, short‑horizon, and rule‑based safety experts. A Gaussian Mixture Model preserves multimodal maneuver branches, improving robustness and preventing behavioral averaging, leading to better safety compliance, decision stability, and generalization in experiments.
PRISM (Proactive Risk Intelligence and Safety Management) is an agentic multi-model architecture designed to shift autonomous transportation safety from reactive crash avoidance to proactive, continuous risk management. It uses inverse crash‑probability modeling to transform binary crash classifiers into dynamic safety scores, and runs three specialized models—trajectory kinematics, environmental risk, and VRU interaction—coordinated by a reinforcement‑learning reasoning layer. Across 1,296 naturalistic driving scenarios, PRISM achieved a mean safety score of 68/100, classified 77.6% of situations as advisory, and flagged 3.8% as near‑misses, with 11% requiring intervention or emergency response, highlighting trajectory risk and VRU proximity as key safety factors.