arXiv:2609.27510v1 Announce Type: cross
Abstract: Modern large language models are pretrained on massive datasets, making it difficult to prevent benchmark data from entering their training sets and...
By Kaifeng Tan, Yudong Li, Linlin Shen
Debias‑SparseGPT is a post‑training pruning technique that adds a representational debiasing term based on demographically contrasting inputs to mitigate bias amplification caused by weight sparsification. The method is validated across various generative LLMs and sparsity levels (25%, 50%, and structured 2:4), consistently reducing pruning‑induced bias while maintaining perplexity and zero‑shot accuracy. In the most aggressive 2:4 sparsity regime, enriching the calibration set with long‑context, content‑rich examples further improves both downstream performance and fairness.
By Irina Proskurina, Guillaume Metzler, Antoine Gourru, Julien Velcin
Manacá-1B is a 1.72‑billion‑parameter, open decoder‑only language model trained from scratch for Brazilian Portuguese, released with a fully containerized, reproducible training pipeline and complete logs. The authors evaluate it against nine open baselines on four Portuguese benchmarks, reporting standard errors and paired significance tests, and find that Manacá-1B outperforms smaller models on LAMBADA‑PT while remaining competitive on commonsense completion. They also uncover a tokenizer‑related evaluation pitfall that can drastically lower accuracy and provide a simple fix, releasing all code, logs, and corrected tokenizer for full reproducibility.
By Bruno Leonardo Santos Menezes, Carlos Leonardo Souza Cardoso, Fabio Andre Machado Porto
arXiv:2609.13154v1 Announce Type: new
Abstract: Recent advances in large language models (LLMs) have made prompts increasingly large and complex. Techniques such as chain-of-thought reasoning (Wei et...
By Shamin Chokshi
Brazilian Portuguese remains under-served by open language models, and the few that exist are difficult to reproduce and are often compared without measures of uncertainty. We release Manacá-1B, an op...
The study evaluates how extractive prompt compressors affect token costs across ten languages, finding that compressors trained on English data widen the token premium gap for non‑English languages, while a multilingual compressor does not. The gap is tied to the supervision data rather than model architecture, and aggressive compression can reduce non‑English contexts to near‑zero utility. A translate‑then‑compress approach can match or outperform native compression at roughly half the token cost in several languages.
By Mantas Lukauskas
As large language models (LLMs) grow more capable, they are increasingly deployed in context-rich settings where task inputs are often accompanied by long, partially irrelevant context. In a controlled setting, we find that state-of-the-art models often appear robust to task-irrelevant context at the aggregate level: prepending it to benchmark questions causes little change in overall accuracy.
The paper introduces DEX-Comp, a two‑stage training method for soft context compression in Retrieval‑Augmented Generation (RAG). First, a pure distillation warm‑start trains the compression model on correct responses from an uncompressed RAG. Then, hard exploration uses reinforcement learning on queries where the uncompressed RAG fails, encouraging better computation patterns for compressed representations. Experiments on five open‑domain QA benchmarks show that DEX‑Comp compresses retrieved contexts 16×, speeds inference 4×–24×, and matches or surpasses the uncompressed RAG baseline across various retrieval depths.
By Shuyu Guo, Shuo Zhang, Zhaochun Ren
arXiv:2606. 24083v1 Announce Type: cross Abstract: "Talk short.
By Morayo Danielle Adeyemi, Ryan A. Rossi, Franck Dernoncourt
arXiv:2609.37076v1 Announce Type: new
Abstract: Large language models trained on vast corpora inherently risk memorizing harmful content that may later re-emerge in their outputs. To mitigate this is...
By Puning Yang, Qizhou Wang, Junchi Yu, Bo Han, Xiuying Chen
arXiv:2606. 12117v1 Announce Type: cross Abstract: Benchmark scores often misrepresent a large language model's (LLM's) knowledge, because they rely, e.
By Selen Erkan, Bastian Boll, Kristian Kersting, Bj\"orn Deiseroth, Letitia Parcalabescu
The paper introduces PILL, a new infilling technique for diffusion language models that eliminates the need for a preset initial length and reduces inference overhead. PILL uses probing-based length-free decoding, cutting down on extra forward passes and speeding up generation. Experiments across five diffusion models and eight benchmarks show PILL outperforms the strongest baseline with higher pass rates and BLEU-2 scores while running 1.82× faster.
By Haobo Xu, Sirui Chen, Yuanchen Bei, Lingjie Chen, Yuchen Yan, Dongqi Fu, Jingrui He, Hanghang Tong