arXiv Computation and Language By Martin Dvo\v{r}\'ak, V\'it Tlusto\v{s}, Artyom Voronin, Martin Habrovec, Kate\v{r}ina Podlesn\'a, Barbora Ri\v{s}ov\'a, Josef Von\'a\v{s}ek

Size Matters: Foundation Model for Czech HTML documents

Read the original on arXiv Computation and Language →

The paper introduces HTML‑LM, a 154‑million‑parameter foundation model designed for Czech HTML documents. It leverages HTML‑aware training and a ModernBERT architecture, trained on 100 million web pages with objectives such as masked language modeling, bag‑of‑words prediction, and contrastive distillation from larger language models. HTML‑LM achieves state‑of‑the‑art performance on classification and regression tasks in the Czech Internet domain, outperforms larger encoders and small LLMs, and is deployed in production to process thousands of web documents per second.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jun 12

SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation

arXiv:2606. 13647v1 Announce Type: cross Abstract: We introduce SkMTEB, the first comprehensive MTEB-style text embedding benchmark for Slovak, a low-resource West Slavic language, comprising 31 datasets across 7 task types -- nearly 4$\times$ the depth of existing multilingual benchmark coverage for Slovak.

By Marek \v{S}uppa, Andrej Ridzik, Daniel Hl\'adek, Nat\'alia K\v{n}a\v{z}ekov\'a, Vikt\'oria Ondrejov\'a
arXiv AI
Aug 13

Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression

arXiv:2608. 11249v1 Announce Type: cross Abstract: We study the problem of lossless text compression, motivated by the rapid growth in the collection and storage of digital textual data - including plain text, source code, and structured formats such as XML - and by recent advances in neural language model-based compression.

By Angelo Nardone, Paolo Ferragina
arXiv AI
Jun 30

PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation

arXiv:2606. 28344v1 Announce Type: cross Abstract: Augmenting large language models (LLMs) with retrieved web text has become a dominant paradigm, yet the web is not natively textual: existing systems depend on complex parsing pipelines that linearize HTML and discard layout, visual structure, and formatting.

By Yichuan Wang, Zhifei Li, Zirui Wang, Paul Teiletche, Lesheng Jin, Matei Zaharia, Joseph E. Gonzalez, Sewon Min