STAR-OPD: Structured Aspect-Cascade-Aware On-Policy Reward Distillation for ABSA Quadruple Extraction
Read the original on arXiv AI →The paper introduces STAR-OPD, a structured aspect‑cascade‑aware on‑policy reward distillation method for aspect‑based sentiment analysis (ABSA) quadruple extraction. It addresses a specific failure mode where distilled models produce structurally invalid target‑aspect bindings, leading to hallucinated targets and corrupted downstream predictions. By training on student rollouts and applying set‑structured rewards that enforce binding consistency, target grounding, and fine‑grained aspect disambiguation, STAR‑OPD outperforms both off‑policy and generic on‑policy baselines on E‑ABSA20K and SemEval‑2014, reducing target hallucination and improving performance on structurally hard cases.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.