arXiv Computation and Language By Se Un Park, Hakjun Kim, Taehoon Roh, Junyoung Park

Personalized Korean Lipreading as Visual Speech Recognition: Transfer, Census and Adaptation on OLKAVS

Read the original on arXiv Computation and Language →

The paper introduces a personalized Korean visual speech recognition system that uses a video-only Conformer model initialized from English-trained weights, achieving a character error rate (CER) of 9.95–12.19% on the OLKAVS nine-camera corpus and 19.00–21.52% on unseen wording. Individual speaker CER varies widely (1.0–52.2%), with seen wording reducing errors by 7.0–9.0 points and professional or spontaneous speech increasing errors by 8.5–12.7 points. A low‑rank adapter, comprising only 4.6% of the model parameters and trained on 4–29 minutes of a user’s frontal video, reduces high‑error speakers’ CER by 2.13–3.58 points, transfers across all cameras without loss, and retains 85% of full fine‑tuning benefits at 12% of its cost; cameras above the mouth plane add a constant offset of about six CER points that can be mitigated by training on all views.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 25

Closing the Quality Gap in Low-Resource Text-to-Speech: LoRA Fine-Tuning of VoxCPM2 for Khmer and Korean

The paper investigates how to close the quality gap in low‑resource text‑to‑speech for Khmer and Korean using the VoxCPM2 model. By training a single low‑rank adaptation (LoRA) adapter on a shared 25.5‑hour corpus, the authors improve Khmer’s mean opinion score from 3.85 to 4.23 with a rank‑64 adapter, while Korean shows no significant gain. The study highlights that adaptation benefits mainly when the base model is weak and that training loss does not always align with human ratings.

By Phannet Pov, Hyun Woo Park, Voneat Pen, Sovandara Chhoun, Wan-Sup Cho, Saksonita Khoeurn