arXiv AI By Swati Rajwal, Sanjay Das, Tirthankar Ghosal

Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis Ranking

Read the original on arXiv AI →

arXiv:2608. 17270v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for scientific hypothesis generation.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.