arXiv AI By Lisette Esp\'in-Noboa, Gonzalo Gabriel M\'endez

Whose Name Comes Up? II: Benchmarking and Intervention-Based Auditing of LLM-Based Scholar Recommendation

Read the original on arXiv AI →

arXiv:2602. 08873v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are now used for academic expert recommendation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 11

MERIT: Matching Expertise via Rubric-Informed Training for Reviewer Assignment

MERIT is a two‑stage framework for reviewer assignment that first trains a reviewer assessor using reinforcement learning to match paper‑specific expertise rubrics with reviewers’ prior work, guided by an LLM judge. The assessor’s predictions are then distilled into an embedding‑based retriever for efficient large‑scale assignment. Experiments show the 4B assessor outperforms larger general‑purpose LLMs on suitability classification, and the retriever achieves state‑of‑the‑art performance on LR‑Bench and the CMU Gold dataset.

By Zixuan Yang, Yibo Zhao, Weicong Liu, Xiang Li
arXiv Computation and Language
Aug 27

IDEAlign: Comparing Ideas of Large Language Models to Domain Expert

IDEAlign introduces a new protocol for evaluating the similarity of large language model (LLM) annotations to expert judgments. It uses pick‑the‑odd‑one‑out tasks to capture expert similarity and benchmarks various similarity methods—including text embeddings, topic models, and LLM-as-a-judge—against these human ratings. Applied to educational datasets, the study finds that most metrics miss nuanced expert dimensions, with LLM-as-a-judge performing best yet still insufficient for full expert alignment.

By Hyunji Nam, Lucia Langlois, James Malamut, Mei Tan, Dorottya Demszky
arXiv AI
Aug 19

Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis Ranking

The paper investigates whether large language models (LLMs) can reliably assess scientific hypotheses by using a logit-based energy scoring method that leverages the model’s intrinsic confidence. Across 1,323 papers in 12 disciplines, this intrinsic scoring achieved a 33.0% Hit@1 rate, outperforming a prompted listwise ranking approach that scored 16.6%. The best result, a 1‑billion‑parameter model with logit-based energy scoring, reached 53.1% Hit@1, suggesting that confidence‑based evaluation could improve trustworthy AI‑enabled scientific discovery.

By Swati Rajwal, Sanjay Das, Tirthankar Ghosal