Hugging Face Blog

vLLM V0 to V1: Correctness Before Corrections in RL

arXiv Machine Learning
Sep 10

Robustness of LLM-Generated SystemVerilog Assertions to Semantics-Preserving RTL Transformations

The paper evaluates how robust large language models are at generating SystemVerilog Assertions (SVA) when the underlying RTL code undergoes semantics‑preserving transformations such as operand reordering, identifier renaming, and redundant parenthesization. Using a curated dataset and two open‑source models (Qwen2.5‑Coder‑7B and DeepSeek‑Coder‑V2‑Lite), the authors find that 9.7%–27.0% of behaviors that were correct on the original RTL become incorrect after transformation, revealing significant instability that aggregate accuracy metrics can hide.

By FNU Aditi
arXiv AI
5d ago

Diagnosing Training Inference Mismatch in LLM Reinforcement Learning via a Zero-Mismatch Reference

The paper investigates Training‑Inference Mismatch (TIM) in large‑language‑model reinforcement learning, where rollout generation and policy optimization produce differing token probabilities despite identical model weights. By creating a zero‑mismatch diagnostic setting called VeXact, the authors isolate TIM and demonstrate that even minor token‑level numerical disagreements can trigger training collapse. They further show that TIM alters the effective optimization problem and propose remedies to mitigate its impact.

By Tianle Zhong, Neiwen Ling, Yifan Pi, Zijun Wei, Tianshu Yu, Geoffrey Fox, Peng Wu, Xiao Yu