Scaling-up BERT Inference on CPU (Part 1)
Related stories
Accelerate BERT inference with Hugging Face Transformers and AWS Inferentia
BERT 101 - State Of The Art NLP Model Explained
Fine-Tune W2V2-Bert for low-resource ASR with 🤗 Transformers
From BERT to Frontier Agents: Eight Years of Language-Model Progress, the Collapse of the Capability-Cost Curve, and the Rise of Task-Targeted Models
arXiv:2608. 13675v1 Announce Type: cross Abstract: Between October 2018 and July 2026 AI models progressed from simple systems like BERT to massive agents that solve complex math and write software.
Looped GPT-BERT: Trading Parameters for Computation in Small Language Modeling
arXiv:2609.09691v1 Announce Type: new Abstract: When training data are limited, increasing parameter count is not the only way to improve language-model performance. A small parameter set, when repea...
Finally, a Replacement for BERT: Introducing ModernBERT
Separating Representation from Reconstruction Enables Scalable Text Encoders
arXiv:2607. 04011v1 Announce Type: cross Abstract: While decoders have rapidly scaled, encoders have remained largely unchanged since BERT.
Scaling laws for neural language models
FOCUS: DLLMs Know How to Tame Their Compute Bound
arXiv:2601. 23278v2 Announce Type: replace Abstract: Diffusion Large Language Models (DLLMs) offer a compelling alternative to Auto-Regressive models, but their deployment is constrained by high decoding cost.
R-DEIM Net: An Efficient Rationale-Augmented Dual-Expert Interaction Model for Paraphrase Detection
R-DEIM Net is a 76‑million‑parameter dual‑expert model designed for paraphrase detection that balances accuracy with computational efficiency. It combines an Interaction Expert, which captures token‑level similarity via multi‑scale 2D convolutions and attention, with a Reasoning Expert that generates human‑readable rationales using a Flan‑T5‑small decoder. On the Quora Question Pairs dataset, the model attains 90.07% accuracy and 90.16% F1‑score, matching strong transformer baselines while producing auxiliary rationales.
BnBERT-iPET: Sparse Few-Shot Language Modeling for Bengali via Lottery Ticket Pruning
arXiv:2608. 05104v1 Announce Type: new Abstract: Deep neural networks have shown impressive success in NLP tasks owing to their complex structure and huge number of edges.