arXiv Computation and Language By Bo Chen

The Public Discourse Corpus (PDC): A Speaker-Attributed Dataset for Valence and Epistemic Modality with Target Speaker Participation

Read the original on arXiv Computation and Language →

The Public Discourse Corpus (PDC) is the first dataset of public‑figure interview speech annotated for affective valence and epistemic modality. It contains 998 videos from 100 speakers across seven professional domains, yielding 186,642 sentences (3.1 million words). A key methodological contribution is Target Speaker Participation (TSP), a five‑category annotation taxonomy with documented inter‑annotator reliability (κ = 0.616), and an audio‑first diarization pipeline that separates target‑speaker turns from interviewer and third‑party speech. The corpus, annotation tools, validation sample, and processing pipeline are released as open source.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.