Mitigating Bus Bunching with Reinforcement Learning Enhanced by Semantic Stop Embedding
arXiv:2608. 10207v1 Announce Type: new Abstract: Bus bunching degrades service regularity and increases passenger waiting in high-frequency transit.
arXiv:2608. 10207v1 Announce Type: new Abstract: Bus bunching degrades service regularity and increases passenger waiting in high-frequency transit.
arXiv:2610.01413v1 Announce Type: cross Abstract: Reinforcement learning (RL) algorithms frequently compare probability distributions, such as state visitation distributions induced by policies and e...
arXiv:2510. 12560v2 Announce Type: replace-cross Abstract: End-to-end autonomous driving models trained with imitation learning (IL) often generalize poorly, particularly in long-tail scenarios where expert demonstrations are sparse.
arXiv:2609.38673v1 Announce Type: new Abstract: Model-based offline reinforcement learning (MBORL) improves sample efficiency through model-generated trajectories. However, accumulative model error c...
arXiv:2510. 01460v4 Announce Type: replace-cross Abstract: Offline-to-online reinforcement learning (RL) has emerged as a practical paradigm that leverages offline datasets for pretraining and online interactions for fine-tuning.
arXiv:2607. 18286v1 Announce Type: cross Abstract: Transit signal priority (TSP) requires balancing competing objectives: reducing bus delay while limiting adverse impacts on non-bus traffic and avoiding extreme waits for a subset of vehicles.
arXiv:2505. 22442v3 Announce Type: replace-cross Abstract: Offline RL (ORL) promises safe and sample-efficient deployment but existing methods rely on undocumented online interactions for hyperparameter tuning and lack reliable fully offline estimates of initial online performance.
The paper presents a compact multimodal policy trained via behavioral cloning to drive autonomously in the CARLA simulator. Using five‑frame histories of RGB images, LiDAR, telemetry, and lane waypoints, the 1.36‑million‑parameter model predicts throttle, brake, and steering at 20 Hz. Trained on 236,882 windows (≈3.3 hours of driving) from 448 captures, the policy drives for hours on both training and unseen routes without collisions, demonstrating qualitative transfer and recovery from large trajectory deviations.
arXiv:2510. 03494v2 Announce Type: replace Abstract: We study finite-horizon offline reinforcement learning (RL) with function approximation for both policy evaluation and policy optimization.
arXiv:2606. 23978v1 Announce Type: cross Abstract: We present an offline reinforcement learning (RL) framework for optimizing SLAM throughput control in a warehouse fulfillment environment.
arXiv:2609.35880v1 Announce Type: new Abstract: Offline reinforcement learning (RL) has traditionally focused on learning policies for direct deployment under conservative objectives, where uncertain...
arXiv:2306. 09712v2 Announce Type: replace Abstract: In reinforcement learning (RL), there are two major settings for interacting with the environment: online and offline.