arXiv Computation and Language

Personalized Korean Lipreading as Visual Speech Recognition: Transfer, Census and Adaptation on OLKAVS

The paper introduces a personalized Korean visual speech recognition system that uses a video-only Conformer model initialized from English-trained weights, achieving a character error rate (CER) of 9.95–12.19% on the OLKAVS nine-camera corpus and 19.00–21.52% on unseen wording. Individual speaker CER varies widely (1.0–52.2%), with seen wording reducing errors by 7.0–9.0 points and professional or spontaneous speech increasing errors by 8.5–12.7 points. A low‑rank adapter, comprising only 4.6% of the model parameters and trained on 4–29 minutes of a user’s frontal video, reduces high‑error speakers’ CER by 2.13–3.58 points, transfers across all cameras without loss, and retains 85% of full fine‑tuning benefits at 12% of its cost; cameras above the mouth plane add a constant offset of about six CER points that can be mitigated by training on all views.

arXiv Computation and Language
Sep 25

Closing the Quality Gap in Low-Resource Text-to-Speech: LoRA Fine-Tuning of VoxCPM2 for Khmer and Korean

The paper investigates how to close the quality gap in low‑resource text‑to‑speech for Khmer and Korean using the VoxCPM2 model. By training a single low‑rank adaptation (LoRA) adapter on a shared 25.5‑hour corpus, the authors improve Khmer’s mean opinion score from 3.85 to 4.23 with a rank‑64 adapter, while Korean shows no significant gain. The study highlights that adaptation benefits mainly when the base model is weak and that training loss does not always align with human ratings.

By Phannet Pov, Hyun Woo Park, Voneat Pen, Sovandara Chhoun, Wan-Sup Cho, Saksonita Khoeurn