arXiv AI By Fabian P. Kr\"uger, Andrea Hunklinger, Adrian Wolny, Tim J. Adler, Igor Tetko, Santiago David Villalba

SEISMO: Explanation-Aware, Trajectory-Conditioned LLM Agents for Sample-Efficient Molecular Optimisation

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Machine Learning
3d ago

Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment

The paper introduces an LLM-as-a-Judge framework for evaluating the outputs of an agentic drug discovery assistant, ChatInvent, deployed at AstraZeneca. It defines four quality dimensions—Completeness, Relevancy, Structural Clarity, and Scope Adherence—alongside deterministic Tool Call Correctness checks, and validates the judge against five expert annotators. After optimizing the best-performing judge with few-shot demonstrations, alignment with human majority votes improves from 0.80 to 0.86, and the framework reveals that informal question phrasing does not degrade output quality.

By Emma Granqvist, Roc\'io Mercado, Samuel Genheden