Flesch-Kincaid Readability Depends Only on the Topic Distribution in Long Texts under Topic Models
Read the original on arXiv Computation and Language →The paper shows that the Flesch Reading Ease and Flesch‑Kincaid Grade Level scores, which are computed from the same two document statistics, converge almost surely to deterministic functions of a document’s topic distribution when modeled with a topic model that includes explicit sentence boundaries. In the long‑text limit, all variation in these scores is driven solely by topical composition, not by any residual readability signal. Experiments on the Brown and BNC corpora demonstrate that a topic vector inferred from one half of a document can predict the other half’s FKGL with substantial correlation (r = 0.779 and 0.884), though adding this prediction to genre and syllable‑count features yields only marginal gains in explained variance.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.