Medium is the new large.
Related stories
Large Enough
Mistral Small 3
What's new in Diffusers? ๐จ
My Tailor is Mistral
Mistral Small 3.1
SmolVLM Grows Smaller โ Introducing the 256M & 500M Models!
DeepSeek V4 Pro 0813 (on OpenRouter)
DeepSeek V4 Pro 0813 (on OpenRouter) The latest DeepSeek Pro model is now available, via API only. I had to link to OpenRouter because DeepSeek don't have any obvious announcement page for their new model.
Diffusers welcomes Stable Diffusion 3.5 Large
Domain-Aware Scaling Laws Uncover Data Synergy
arXiv:2607. 11052v1 Announce Type: new Abstract: Machine learning progress is often attributed to scaling model size and dataset volume, yet the composition of data can be just as consequential.
$\mu$pscaling small models: Principled warm starts and hyperparameter transfer
arXiv:2602. 10545v2 Announce Type: replace-cross Abstract: Modern large-scale neural networks are often trained and released in multiple sizes to accommodate diverse inference budgets.
Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts
arXiv:2608. 20061v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures significantly expand model capacity without a proportional increase in computational cost.