ESTS at WMT26: Routing-Informed Expert Pruning for Model Compression
Read the original on arXiv Machine Learning →The paper reports six submissions by the ESTS team to the WMT26 Model Compression Shared Task for English–Simplified Chinese and English–Egyptian Arabic. Each submission offers three compression operating points derived from GPT‑OSS‑20B, using routing‑informed expert pruning, cross‑lingual routing divergence for capacity allocation, and MXFP4 quantization of retained expert projection weights. The resulting models, ranging from 4.186 B to 7.770 B parameters, are fine‑tuned on GPT‑5.1 synthetic data and evaluated internally with xCOMET‑XL.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.