arXiv AI
Sep 18

SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science

SCICONVBENCH is a benchmark designed to evaluate large language models (LLMs) on multi‑turn clarification tasks in computational science. It focuses on two key abilities: eliciting missing information (disambiguation) and resolving contradictory requests (inconsistency resolution) across four domains—fluid mechanics, solid mechanics, materials science, and partial differential equations. The benchmark pairs a structured task ontology with a rubric‑based evaluation framework, measuring LLM performance in clarification behavior, conversational grounding, and final‑specification fidelity, and reveals that even top models only resolve about 52.7% of disambiguation cases in fluid mechanics while often making ungrounded assumptions.

By Nithin Somasekharan, Youssef Hassan, Shiyao Lin, Gihan Panapitiya, Patrick Emami, Anurag Acharya, Sameera Horawalavithana, Shaowu Pan