arXiv Computation and Language
Sep 2

PersianAnonymizer: Evaluating LLM-Labeled Training for Efficient NER-based Anonymization in Persian

The paper presents PersianAnonymizer, a method for anonymizing Persian customer chats by training a compact NER model using supervision from large language models (LLMs). Three instruction‑tuned LLMs—DeepSeek‑V3‑0324, GPT‑OSS‑120B, and Qwen3‑235B‑A22B‑Instruct‑2507—were used to generate span annotations, producing four corpora. A MatinaRoberta token‑classifier trained on each corpus achieved high macro‑F1 and Label Coverage Recall, with the OSS_ZeroShot‑derived NER labeling a 40K‑message test set in about two minutes on a single RTX 3090, demonstrating a practical, low‑cost approach to Persian data anonymization.

By Mohammad Hossein Shalchian, Mostafa Amiri, Amir Mahdi Sadeghzadeh
arXiv Machine Learning
Jun 5

Domain-Adapted Small Language Models with Hybrid Post-Processing: Achieving Cost-Efficient, Low-Latency Multi-Label Structured Prediction via LoRA Fine-Tuning on Scarce Data

arXiv:2606. 05781v1 Announce Type: new Abstract: Deploying frontier large language models (LLMs) for domain-specific structured evaluation tasks often incurs substantial latency, cost, and data privacy overhead.

By Srinivasan Manoharan, Dilipkumar Nallusamy, Sachin Kumar, Haifeng Wu
arXiv Computation and Language
Sep 1

Manac\'a-1B: An Open, Reproducible Brazilian-Portuguese Language Model and a Tokenizer-Aware, Paired Evaluation

Manacá-1B is a 1.72‑billion‑parameter, open decoder‑only language model trained from scratch for Brazilian Portuguese, released with a fully containerized, reproducible training pipeline and complete logs. The authors evaluate it against nine open baselines on four Portuguese benchmarks, reporting standard errors and paired significance tests, and find that Manacá-1B outperforms smaller models on LAMBADA‑PT while remaining competitive on commonsense completion. They also uncover a tokenizer‑related evaluation pitfall that can drastically lower accuracy and provide a simple fix, releasing all code, logs, and corrected tokenizer for full reproducibility.

By Bruno Leonardo Santos Menezes, Carlos Leonardo Souza Cardoso, Fabio Andre Machado Porto