arXiv AI By Alyssa Unell, Natalie Dullerud, Naomi Boneh, Meena Jagadeesan, Tatsu Hashimoto, Nigam Shah, Sanmi Koyejo

Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability

Read the original on arXiv AI →

arXiv:2606. 15029v1 Announce Type: new Abstract: LLM judges are used to reduce the need for costly human labor in evaluating open-ended text generation.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Jul 28

Designing Service Systems from Textual Evidence

arXiv:2603. 10400v2 Announce Type: replace-cross Abstract: Designing service systems requires selecting among alternative configurations -- choosing the best chatbot variant, the optimal routing policy, or the most effective quality control procedure.

By Ruicheng Ao, Hongyu Chen, Siyang Gao, Hanwei Li, David Simchi-Levi