arXiv AI

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition

arXiv:2607. 13347v2 Announce Type: replace-cross Abstract: LLM-as-a-judge is widely used to provide feedback and selection signals in closedloop regeneration, but this use remains insufficiently validated.