The article reviews how transformer models, retrieval‑grounded pipelines, and large language models are reshaping hadith computational science. It critically evaluates existing literature, highlighting uneven progress: expanded data resources and mature segmentation tasks, yet persistent issues such as narrow corpora, weak benchmark comparability, and limited reproducibility. The authors argue that hadith computation should be viewed as an evidence‑infrastructure problem requiring knowledge integration, provenance, and expert supervision, and they propose a research agenda to strengthen the field’s methodological rigor.
By Md. Ashraful Haque (Greentech Apps Foundation, United Kingdom), Riasat Islam (Greentech Apps Foundation, United Kingdom, Queen Mary University of London, London, United Kingdom)
Ansari is a retrieval‑grounded Islamic AI assistant that has handled over 140,000 conversations in more than 25 languages since June 2023. It uses an agentic retrieval loop where a language model searches authenticated Islamic corpora—including the Qur’an, hadith collections, fiqh encyclopedias, and tafsir sources—and answers only based on retrieved content, providing citations for verification. The paper details Ansari’s architecture, multi‑platform deployment, evaluation results (including top performance on the IslamicMMLU leaderboard and strong resistance to false premises), and lessons for faith‑sensitive LLM deployments.
By M Waleed Kadous, Amr Elsayed, Abdullah Al Nahas, Ashraf Haress
TruthInsightBench is a new benchmark designed to evaluate automated scientific discovery agents by presenting them with 40 blind tasks drawn from peer‑reviewed studies across ten domains. Each task provides only a neutral objective and frozen data, withholding source conclusions, expected values, and analysis paths, forcing agents to determine which claim the data support. A fixed LLM‑based judge scores agents on evidentiary maturity across six dimensions, using 29 artifact‑grounded items, enabling fully automated, repeatable evaluation without human grading.
By Zhibo Yang, Chen Zhang, Yuewei Zhang, Hao Wang
arXiv:2609.26035v1 Announce Type: new
Abstract: Conversational agents often express answers in a uniformly confident register. We test whether expressed uncertainty, provenance-aware assertion, and e...
By Sebastian Cochinescu
arXiv:2606. 04990v2 Announce Type: replace-cross Abstract: Large language model (LLM)-based agents are evolving from passive text generators into autonomous systems capable of planning, tool use, retrieval, memory access, environmental interaction, and multi-agent collaboration.
By Yiqi Wang, Jiaqi Zhang, Taotao Cai, Zirui Liu, Qingqiang Sun, Zequn Sun, Zhangkai Wu, Manqing Dong, Mingkai Zhang, Xuefei Yin, Yanming Zhu
arXiv:2606. 04990v1 Announce Type: cross Abstract: Large language model (LLM)-based agents increasingly solve complex tasks by interacting with external tools, retrieval systems, memory modules, environments, and other agents.
By Yiqi Wang, Jiaqi Zhang, Taotao Cai, Zirui Liu, Qingqiang Sun, Zequn Sun, Zhangkai Wu, Mingkai Zhang, Yanming Zhu
VeriPhy is an auditable physical‑verification system that evaluates generated video by compiling prompts into typed physical obligations and a statically validated execution plan before any frames are observed. During execution, it gates calls to frozen low‑level experts (e.g., segmentation, tracking, counting, depth, OCR, audio‑event detection) and returns provenance‑carrying evidence records, which are mapped to a three‑valued state (supported, contradicted, unknown) with full traceability. On a 1,500‑clip corpus of human‑annotated flaw records, VeriPhy accounts for 228 failures out of 304, outperforming a published question‑decomposition evaluator that accounts for 164, while also providing auditable evidence for each verdict.
arXiv:2606. 18037v1 Announce Type: new Abstract: Tool-using LLM agents increasingly use the Model Context Protocol (MCP) to answer from heterogeneous evidence sources, including search, APIs, databases, clinical records, and formulary tools.
By Ander Alvarez, Santhiya Rajan, Samuel Mugel, Rom\'an Or\'us
arXiv:2607. 01223v1 Announce Type: new Abstract: When should an AI system's answer be trusted?
By Ben Slivinski, Michael Saldivar
VeriPhy is an auditable physical‑verification system that transforms a text prompt into typed physical obligations and a statically validated execution plan before any video frames are generated. During execution, it gates calls to frozen low‑level experts (segmentation, tracking, counting, depth, OCR, audio‑event detection, etc.) and records provenance‑carrying evidence for each action. The system maps these records to a three‑valued state—supported, contradicted, or unknown—providing traceable verdicts that can be used to refine generation models.
By Wenzhuo Xu, Yuchen Zhu, Chongjian Ge, Xuan Shen, Jing Shi, Jason Kuen, Yongxin Chen, Molei Tao, Christopher McComb, Noelia Grande Guti\'errez, Jiuxiang Gu
arXiv:2607. 18240v1 Announce Type: new Abstract: Large language models (LLMs) can achieve strong fact-checking accuracy, yet forced binary decisions conceal a critical reliability problem: systems may issue confident verdicts even when supporting evidence is weak, sparse, or internally inconsistent.
By Dekun Yang
arXiv:2607. 17883v1 Announce Type: cross Abstract: Enterprises will not deploy AI agents they cannot trust, and the most-cited reason for distrust is hallucination: confident, fluent output that is simply not true.
By Bogdan Raduta, Horia Velicu, Alexandru Preda, Serban Chiricescu