Personalized Korean Lipreading as Visual Speech Recognition: Transfer, Census and Adaptation on OLKAVS
Read the original on arXiv Computation and Language →The paper introduces a personalized Korean visual speech recognition system that uses a video-only Conformer model initialized from English-trained weights, achieving a character error rate (CER) of 9.95–12.19% on the OLKAVS nine-camera corpus and 19.00–21.52% on unseen wording. Individual speaker CER varies widely (1.0–52.2%), with seen wording reducing errors by 7.0–9.0 points and professional or spontaneous speech increasing errors by 8.5–12.7 points. A low‑rank adapter, comprising only 4.6% of the model parameters and trained on 4–29 minutes of a user’s frontal video, reduces high‑error speakers’ CER by 2.13–3.58 points, transfers across all cameras without loss, and retains 85% of full fine‑tuning benefits at 12% of its cost; cameras above the mouth plane add a constant offset of about six CER points that can be mitigated by training on all views.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.