vLLM V0 to V1: Correctness Before Corrections in RL
Related stories
Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR
arXiv:2605. 02909v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become a powerful approach for improving the reasoning capabilities of large language models (LLMs).
CASS-RTL: Correctness-Aware Subspace Steering for RTL Generation with LLMs
arXiv:2606. 05680v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have enabled the automatic synthesis (generation) of register-transfer level (RTL) code from natural language instructions, offering a promising pathway to accelerate chip design.
Gotta Learn Fast: A new benchmark for generalization in RL
Open R1: Update #2
Learning Self-Correction in Vision-Language Models via Rollout Augmentation
arXiv:2602. 08503v2 Announce Type: replace-cross Abstract: Self-correction is essential for solving complex reasoning problems in vision-language models (VLMs).
Geometry-Preserving Orthonormal Initialization for Low-Rank Adaptation in RLVR
arXiv:2606. 31813v1 Announce Type: cross Abstract: Low-rank adaptation (LoRA) and its variants enable parameter-efficient fine-tuning of large language models under the supervised fine-tuning (SFT) paradigm.
LLM4RTL: Tool-Assisted LLM for RTL Generation
arXiv:2606. 15500v1 Announce Type: cross Abstract: Large language models (LLMs) have facilitated impressive progress in software engineering, code generation, tooling, and systems.
Open R1: Update #4
Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs
arXiv:2608. 11573v1 Announce Type: cross Abstract: Achieving effective self-correction, where models verify and correct their own mistakes, remains a fundamental challenge for large language models (LLMs).
Open-Source LLM-Driven Formal Verification: A Multi-Agent Pipeline for RTL Repair
arXiv:2607. 28877v1 Announce Type: cross Abstract: Verification consumes the majority of modern chip design effort, yet the formal verification tools that provide mathematical guarantees of correctness remain expensive and restrictively licensed.