arXiv Computation and Language
Sep 3

Debias-SparseGPT: Bias-Aware Pruning for Large Language Models

Debias‑SparseGPT is a post‑training pruning technique that adds a representational debiasing term based on demographically contrasting inputs to mitigate bias amplification caused by weight sparsification. The method is validated across various generative LLMs and sparsity levels (25%, 50%, and structured 2:4), consistently reducing pruning‑induced bias while maintaining perplexity and zero‑shot accuracy. In the most aggressive 2:4 sparsity regime, enriching the calibration set with long‑context, content‑rich examples further improves both downstream performance and fairness.

By Irina Proskurina, Guillaume Metzler, Antoine Gourru, Julien Velcin
arXiv Computation and Language
Sep 1

Manac\'a-1B: An Open, Reproducible Brazilian-Portuguese Language Model and a Tokenizer-Aware, Paired Evaluation

Manacá-1B is a 1.72‑billion‑parameter, open decoder‑only language model trained from scratch for Brazilian Portuguese, released with a fully containerized, reproducible training pipeline and complete logs. The authors evaluate it against nine open baselines on four Portuguese benchmarks, reporting standard errors and paired significance tests, and find that Manacá-1B outperforms smaller models on LAMBADA‑PT while remaining competitive on commonsense completion. They also uncover a tokenizer‑related evaluation pitfall that can drastically lower accuracy and provide a simple fix, releasing all code, logs, and corrected tokenizer for full reproducibility.

By Bruno Leonardo Santos Menezes, Carlos Leonardo Souza Cardoso, Fabio Andre Machado Porto
arXiv AI
Aug 28

Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors

The study evaluates how extractive prompt compressors affect token costs across ten languages, finding that compressors trained on English data widen the token premium gap for non‑English languages, while a multilingual compressor does not. The gap is tied to the supervision data rather than model architecture, and aggressive compression can reduce non‑English contexts to near‑zero utility. A translate‑then‑compress approach can match or outperform native compression at roughly half the token cost in several languages.

By Mantas Lukauskas