Do Cantonese-Adapted Language Models Better Predict Cantonese Reading? A Cross-Model Eye-Tracking Evaluation
Read the original on arXiv Computation and Language →The study investigates whether language models trained specifically on Cantonese better predict human reading patterns by comparing eye-tracking data with information-theoretic metrics derived from several models. Two adaptation contrasts were examined: a lightly adapted CKIP GPT-2 Tiny versus its Cantonese derivative JED351, and a heavily adapted Qwen2.5-7B versus CantoneseLLM-7B. Results show that the extensively Cantonese-trained CantoneseLLM-7B consistently outperforms others on lexical surprisal and joint metrics, while entropy reduction favors the less adapted CKIP model, indicating that deeper Cantonese-specific training can improve predictive alignment but that rankings vary by metric.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.