Mini-R1: Reproduce Deepseek R1 „aha moment“ a RL tutorial
Related stories
From DeepSeek V3 to V3.2: Architecture, Sparse Attention, and RL Updates
Understanding How DeepSeek's Flagship Open-Weight Models Evolved
One Year Since the “DeepSeek Moment”
Deep R Programming
arXiv:2301. 01188v5 Announce Type: replace-cross Abstract: Deep R Programming is a comprehensive and in-depth introductory course on one of the most popular languages for data science.
Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design
The paper examines how evaluation design can cause significant fluctuations in the benchmark results of reasoning models, particularly the Deepseek‑R1‑Distill series. It shows that subtle changes in evaluation conditions lead to large variations in reported performance, a phenomenon also seen in other open‑source models fine‑tuned from Deepseek‑R1‑Distill and in the QwQ‑32B model. The authors call for a more rigorous evaluation paradigm and provide empirical assessments of the Deepseek‑R1‑Distill models.
The State Of LLMs 2025: Progress, Problems, and Predictions
A 2025 review of large language models, from DeepSeek R1 and RLVR to inference-time scaling, benchmarks, architectures, and predictions for 2026.
Introducing RWKV - An RNN with the advantages of a transformer
DeepSeek V4 Pro 0813 (on OpenRouter)
DeepSeek V4 Pro 0813 (on OpenRouter) The latest DeepSeek Pro model is now available, via API only. I had to link to OpenRouter because DeepSeek don't have any obvious announcement page for their new model.
Gotta Learn Fast: A new benchmark for generalization in RL
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
arXiv:2606. 19348v1 Announce Type: cross Abstract: We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.
DeepSeek-V4: a million-token context that agents can actually use
The Active Ingredient in Muon's Grokking
arXiv:2607. 20512v1 Announce Type: cross Abstract: The Muon optimizer reaches the grokking threshold on modular arithmetic faster than AdamW.

