How to Choose Between Small and Frontier Models
The rise of small language models The post How to Choose Between Small and Frontier Models appeared first on Towards Data Science .
Still a long way to go, but the future is promising The post Setting Up Your Own Large Language Model appeared first on Towards Data Science .
The rise of small language models The post How to Choose Between Small and Frontier Models appeared first on Towards Data Science .
The paper surveys language models created for Portuguese, noting that while rapid progress has been made in NLP, development has been uneven across languages. It systematically maps 46 Portuguese models, detailing aspects such as base model, architecture, resources, datasets, licensing, code, data, and weights. The study also traces model evolution phylogenetically, highlights research gaps, and outlines future directions for Portuguese language modeling.
The article explains how to transform a small open‑source Qwen LLM into a fast, single‑pass text classifier by replacing its language‑modeling head with a JEV model. It provides a step‑by‑step guide to swapping the head, enabling the LLM to perform classification tasks efficiently. The process leverages the flexibility of open‑source models to create a lightweight, high‑performance classifier.
Cohere, OpenAI, and AI21 Labs have developed a preliminary set of best practices applicable to any organization developing or deploying large language models.
The paper introduces HTML‑LM, a 154‑million‑parameter foundation model designed for Czech HTML documents. It leverages HTML‑aware training and a ModernBERT architecture, trained on 100 million web pages with objectives such as masked language modeling, bag‑of‑words prediction, and contrastive distillation from larger language models. HTML‑LM achieves state‑of‑the‑art performance on classification and regression tasks in the Czech Internet domain, outperforms larger encoders and small LLMs, and is deployed in production to process thousands of web documents per second.
Our latest research finds we can improve language model behavior with respect to specific behavioral values by fine-tuning on a small, curated dataset.