Scaling-up BERT Inference on CPU (Part 1)
Related stories
Accelerate BERT inference with Hugging Face Transformers and AWS Inferentia
BERT 101 - State Of The Art NLP Model Explained
Fine-Tune W2V2-Bert for low-resource ASR with ๐ค Transformers
From BERT to Frontier Agents: Eight Years of Language-Model Progress, the Collapse of the Capability-Cost Curve, and the Rise of Task-Targeted Models
arXiv:2608. 13675v1 Announce Type: cross Abstract: Between October 2018 and July 2026 AI models progressed from simple systems like BERT to massive agents that solve complex math and write software.
Finally, a Replacement for BERT: Introducing ModernBERT
Separating Representation from Reconstruction Enables Scalable Text Encoders
arXiv:2607. 04011v1 Announce Type: cross Abstract: While decoders have rapidly scaled, encoders have remained largely unchanged since BERT.
Scaling laws for neural language models
FOCUS: DLLMs Know How to Tame Their Compute Bound
arXiv:2601. 23278v2 Announce Type: replace Abstract: Diffusion Large Language Models (DLLMs) offer a compelling alternative to Auto-Regressive models, but their deployment is constrained by high decoding cost.
BnBERT-iPET: Sparse Few-Shot Language Modeling for Bengali via Lottery Ticket Pruning
arXiv:2608. 05104v1 Announce Type: new Abstract: Deep neural networks have shown impressive success in NLP tasks owing to their complex structure and huge number of edges.
Ulysses Sequence Parallelism: Training with Million-Token Contexts
Energy-Efficient GPU DVFS for Fine-Tuning of SLMs on Resource-constrained Embedded Devices
arXiv:2607. 05933v1 Announce Type: cross Abstract: Dynamic Voltage Frequency Scaling (DVFS) on resource-constrained embedded GPU platforms is essential for energy-efficient small language model (SLM) fine-tuning, as privacy- and personalization-driven adaptation increasingly requires local execution and involves repeated forward-backward optimization over many mini-batches, making it substantially more time- and energy-intensive than single-pass inference.