Scaling up BERT-like model Inference on modern CPU - Part 2
Related stories
Accelerate BERT inference with Hugging Face Transformers and AWS Inferentia
BERT 101 - State Of The Art NLP Model Explained
Scaling laws for neural language models
From BERT to Frontier Agents: Eight Years of Language-Model Progress, the Collapse of the Capability-Cost Curve, and the Rise of Task-Targeted Models
arXiv:2608. 13675v1 Announce Type: cross Abstract: Between October 2018 and July 2026 AI models progressed from simple systems like BERT to massive agents that solve complex math and write software.
Fine-Tune W2V2-Bert for low-resource ASR with 🤗 Transformers
FOCUS: DLLMs Know How to Tame Their Compute Bound
arXiv:2601. 23278v2 Announce Type: replace Abstract: Diffusion Large Language Models (DLLMs) offer a compelling alternative to Auto-Regressive models, but their deployment is constrained by high decoding cost.
Finally, a Replacement for BERT: Introducing ModernBERT
Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory
Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independently. Memory Decoder introduces a parametric long-term memory module but only studies it at a relatively small scale.
Separating Representation from Reconstruction Enables Scalable Text Encoders
arXiv:2607. 04011v1 Announce Type: cross Abstract: While decoders have rapidly scaled, encoders have remained largely unchanged since BERT.
The State Of LLMs 2025: Progress, Problems, and Predictions
A 2025 review of large language models, from DeepSeek R1 and RLVR to inference-time scaling, benchmarks, architectures, and predictions for 2026.
$\mu$pscaling small models: Principled warm starts and hyperparameter transfer
arXiv:2602. 10545v2 Announce Type: replace-cross Abstract: Modern large-scale neural networks are often trained and released in multiple sizes to accommodate diverse inference budgets.
