arXiv:2608. 05982v1 Announce Type: new Abstract: Inadequate target--disease linkage accounts for 40--50\% of Phase~II efficacy failures, so anticipating which programmes will advance would let sponsors back the hypotheses most likely to reach patients.
By Pui Chung Siu, Claudia Cabrera, Mani Mudaliar, Arkaitz Zubiaga
The paper introduces HypoKG, a unified biochemical knowledge graph built from KEGG, Rhea, and UniProt, and uses it to benchmark 13,200 biomedical hypotheses generated by six large language models (LLMs). By varying the biological information provided—source enzyme only, full biological path, or source and disease endpoint—the study finds that LLMs produce higher-scoring hypotheses when given minimal information, but these are less evidence‑grounded. When supplied with the full biological path, the models generate hypotheses that align more closely with known mechanistic relationships, a phenomenon the authors term evidence‑disciplined reasoning, which is confirmed by shuffling intermediate path steps.
"whyItMatters":"The study demonstrates that knowledge graphs can both uncover novel disease–enzyme pairs and guide LLMs to reason more accurately from evidence, improving the reliability of AI‑generated biomedical hypotheses."
By Dominic Okonkwo, Adetayo Okunoye, Ismailcem Budak Arpinar
arXiv:2606. 01042v1 Announce Type: cross Abstract: Perturbation experiments are central to understanding cellular mechanisms, but remain costly and sparse, motivating prediction of gene expression responses for unobserved conditions.
By Xinyu Yuan, Xixian Liu, Jianan Zhao, Yashi Zhang, Hongyu Guo, Jian Tang
arXiv:2609.06779v1 Announce Type: cross
Abstract: Drug repurposing aims to identify new therapeutic uses for existing compounds and, compared with de novo drug discovery, offers a faster and more cos...
By Zijie Liu, Hongxuan Li, Zhen Tan, Jinhao Duan, Baixiang Huang, Zunpeng Liu, Kai Shu, Tianlong Chen
arXiv:2607. 20163v1 Announce Type: cross Abstract: The rapid growth of biomedical knowledge has made the validation of automatically generated biological annotations a major bottleneck in biomedical curation.
By Emanuele Cavalleri, Miad Alavinezhad, Dario Malchiodi, Marco Mesiti
OptimusKG is a multimodal biomedical labeled property graph that integrates structured and semi‑structured resources to preserve detailed, type‑specific metadata across molecular, anatomical, clinical, and environmental domains. The graph contains nearly 191,000 nodes, over 21.8 million edges, and more than 67 million property instances derived from 18 ontologies, with a top‑level schema that enforces node and edge constraints while retaining granular provenance. Validation using the PaperQA3 agent found that 70.0% of sampled edges are supported by literature evidence, and the graph offers a standardized resource for machine learning, knowledge‑grounded retrieval, and hypothesis generation in biomedical research.
By Lucas Vittor, Ayush Noori, I\~naki Arango, Joaqu\'in Polonuer, Sam Rodriques, Andrew White, David A. Clifton, Marinka Zitnik
EvidenceNet is a disease‑specific dataset that transforms full‑text biomedical literature into structured evidence records and graph representations, preserving study design, provenance, and quantitative support. Using an LLM‑assisted pipeline, it extracts experimentally grounded findings, normalizes entities, scores evidence quality, and links related records via typed semantic relations. The released subsets—EvidenceNet‑HCC and EvidenceNet‑CRC—contain thousands of evidence records and richly connected graphs, with high extraction and relation‑type accuracy, enabling retrieval‑augmented question answering and graph‑based tasks such as link prediction and target prioritization.
By Chang Zong, Jinyu Chen, Sicheng Lv, Si-tu Xue, Huilin Zheng, Jian Wan, Lei Zhang
arXiv:2603. 03322v2 Announce Type: replace-cross Abstract: Recent advancements in Large Language Model (LLM) agents have demonstrated remarkable potential in automatic knowledge discovery.
By Chaoqun Yang, Xinyu Lin, Shulin Li, Wenjie Wang, Ruihan Guo, Fuli Feng, Tat-Seng Chua
The paper introduces the Relational Hypergraph Transformer (RHT), a unified architecture that models relational databases as hypergraphs and learns pentadimensional embeddings (PentE). RHT applies sparse relational attention whose complexity scales with the average relational degree, making it computationally efficient for large, high‑dimensional, and high‑cardinality datasets. Experiments on the Synthea synthetic electronic health record dataset show that RHT produces more semantically coherent embeddings than tabular, relational, and temporal graph baselines, while remaining scalable, and the authors provide an open‑source implementation and plan clinical validation on MIMIC‑IV.
By Edouard Lansiaux, Hugo Kazzi, Aur\'elien Loison, Slim Hammadi, Emmanuel Chazard
The paper introduces a time‑aligned evolving concept graph framework that jointly models semantic and structural changes in scientific literature. By treating dated papers as shared update events, it reconstructs both semantic and structural states from the same publication history for each prediction time, and fuses these states at the pair level to forecast co‑occurrence, relation formation, and conditional relation type. Experiments on a large graph of 187,848 papers and 270,687 concepts show that refreshing context with graph updates boosts mean relation AUPRC by 16.6% and raises mean relation AUROC from 0.9290 to 0.9722.
By Fred Sun, Jingze Wang, Minkun Xu, Shangqi Guo
arXiv:2606. 31171v1 Announce Type: new Abstract: Acquiring comprehensive cross-domain biomedical profiles is often costly and time-consuming, resulting in severe data scarcity in medical research.
By Mengying Zhou, Yongjie Yin, Haoyan Xin, Guoping Liu, Yang Chen
arXiv:2608. 06253v1 Announce Type: new Abstract: Metabolomics knowledge is distributed across heterogeneous resources and remains difficult to translate into predictive representations.
By Dohyun Ku, Min Gu Kwak, Francisco J. Pasquel, Jing Li