arXiv AI By Jhessica Silva, Carlos Caetano, Helena Maia, Breno Bernard Nicolau de Fran\c{c}a, Sandra Avila, Helio Pedrini

Language Models for Portuguese: A Systematic Mapping Study

Read the original on arXiv AI →

The paper surveys language models created for Portuguese, noting that while rapid progress has been made in NLP, development has been uneven across languages. It systematically maps 46 Portuguese models, detailing aspects such as base model, architecture, resources, datasets, licensing, code, data, and weights. The study also traces model evolution phylogenetically, highlights research gaps, and outlines future directions for Portuguese language modeling.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 1

Manac\'a-1B: An Open, Reproducible Brazilian-Portuguese Language Model and a Tokenizer-Aware, Paired Evaluation

Manacá-1B is a 1.72‑billion‑parameter, open decoder‑only language model trained from scratch for Brazilian Portuguese, released with a fully containerized, reproducible training pipeline and complete logs. The authors evaluate it against nine open baselines on four Portuguese benchmarks, reporting standard errors and paired significance tests, and find that Manacá-1B outperforms smaller models on LAMBADA‑PT while remaining competitive on commonsense completion. They also uncover a tokenizer‑related evaluation pitfall that can drastically lower accuracy and provide a simple fix, releasing all code, logs, and corrected tokenizer for full reproducibility.

By Bruno Leonardo Santos Menezes, Carlos Leonardo Souza Cardoso, Fabio Andre Machado Porto
arXiv Computation and Language
Sep 4

Language, Language Models, and What We're Talking About

The paper discusses how language models, often treated as technical artifacts, are actually shaped by the linguistic data used in their training. Using Italian language models trained on translated and synthetic data, the author questions whether these models truly represent Italian or language more broadly, and whether NLP should focus on producing natural language. The work calls for a clearer distinction between models built as products and those built as tools for linguistic study, suggesting that diverse answers and languages may emerge without necessarily being pessimistic.

By Malvina Nissim