Strong Multilingual Privacy Tagging at Encoder Speed
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The paper presents PersianAnonymizer, a method for anonymizing Persian customer chats by training a compact NER model using supervision from large language models (LLMs). Three instruction‑tuned LLMs—DeepSeek‑V3‑0324, GPT‑OSS‑120B, and Qwen3‑235B‑A22B‑Instruct‑2507—were used to generate span annotations, producing four corpora. A MatinaRoberta token‑classifier trained on each corpus achieved high macro‑F1 and Label Coverage Recall, with the OSS_ZeroShot‑derived NER labeling a 40K‑message test set in about two minutes on a single RTX 3090, demonstrating a practical, low‑cost approach to Persian data anonymization.
arXiv:2608. 19957v1 Announce Type: new Abstract: Natural language code retrieval is a rapidly evolving task in computer science.
arXiv:2606. 05781v1 Announce Type: new Abstract: Deploying frontier large language models (LLMs) for domain-specific structured evaluation tasks often incurs substantial latency, cost, and data privacy overhead.
arXiv:2608. 02616v2 Announce Type: replace-cross Abstract: We present what is, to our knowledge, the first systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.
Manacá-1B is a 1.72‑billion‑parameter, open decoder‑only language model trained from scratch for Brazilian Portuguese, released with a fully containerized, reproducible training pipeline and complete logs. The authors evaluate it against nine open baselines on four Portuguese benchmarks, reporting standard errors and paired significance tests, and find that Manacá-1B outperforms smaller models on LAMBADA‑PT while remaining competitive on commonsense completion. They also uncover a tokenizer‑related evaluation pitfall that can drastically lower accuracy and provide a simple fix, releasing all code, logs, and corrected tokenizer for full reproducibility.
arXiv:2607. 02079v1 Announce Type: cross Abstract: We present HaloGuard 1.