Size Matters: Foundation Model for Czech HTML documents
Read the original on arXiv Computation and Language →The paper introduces HTML‑LM, a 154‑million‑parameter foundation model designed for Czech HTML documents. It leverages HTML‑aware training and a ModernBERT architecture, trained on 100 million web pages with objectives such as masked language modeling, bag‑of‑words prediction, and contrastive distillation from larger language models. HTML‑LM achieves state‑of‑the‑art performance on classification and regression tasks in the Czech Internet domain, outperforms larger encoders and small LLMs, and is deployed in production to process thousands of web documents per second.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.