arXiv Machine Learning By Hongmin Li

The Calibration Turn in AI-Assisted Research: A Conceptual and Methodological Framework for Evidence-Licensed Claims

Read the original on arXiv Machine Learning →

arXiv:2606. 31273v1 Announce Type: new Abstract: AI-assisted research has entered a stage in which the central question is not only whether systems can generate hypotheses, run experiments, or produce manuscripts, but whether their scientific claims are calibrated to the evidence that supports them.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 18

Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness

arXiv:2606. 18874v1 Announce Type: new Abstract: AI systems can increasingly automate scientific workflows, but the reasoning that links prior evidence, generated ideas, experiments and final claims often remains implicit inside model inference.

By Zijian Wang, Hanqi Li, Ziyue Yang, Zijian Hu, Shenghan Zuo, Yunzhe Zhang, Da Ma, Danyu Luo, Chenrun Wang, Jing Peng, Tiancheng Huang, Sijia Guo, Huayang Wang, Zichen Zhu, Senyu Han, Yilu Cao, Kai Yu, Lu Chen
arXiv AI
Aug 26

Ethical LLM-Assisted Research: A Framework for Responsible Delegation, Verification, and Epistemic Value

The paper proposes a normative framework for ethical use of large language models (LLMs) in scientific research, treating reasoning as a distributed process where human control remains essential for epistemic legitimacy. It introduces key constructs—content origin, human verification, responsibility assignment, accountable ownership, and epistemic outcome—to separate claim provenance from verification and responsibility. The authors argue that the ethical boundary hinges on adequate verification and accountable human ownership, and they propose an "epistemic audit" to document delegation, verification, provenance, and responsibility for transparent, reviewable AI-assisted reasoning.

By Kalin Stoyanov
arXiv AI
Sep 7

TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents

TruthInsightBench is a new benchmark designed to evaluate automated scientific discovery agents by presenting them with 40 blind tasks drawn from peer‑reviewed studies across ten domains. Each task provides only a neutral objective and frozen data, withholding source conclusions, expected values, and analysis paths, forcing agents to determine which claim the data support. A fixed LLM‑based judge scores agents on evidentiary maturity across six dimensions, using 29 artifact‑grounded items, enabling fully automated, repeatable evaluation without human grading.

By Zhibo Yang, Chen Zhang, Yuewei Zhang, Hao Wang