Incorporating LLM Embeddings for Variation Across the Human Genome
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2602. 17162v3 Announce Type: replace Abstract: Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the "Laws of Nature".
arXiv:2511. 09026v2 Announce Type: replace-cross Abstract: Whole-genome sequencing (WGS) has revealed numerous non-coding short variants whose functional impacts remain poorly understood.
The paper introduces a comprehensive privacy evaluation framework for genomic language models (GLMs) that quantifies memorization risks using perplexity-based detection, canary sequence extraction, and membership inference. By planting canary sequences at different repetition rates in synthetic and real datasets, the authors systematically assess how repetition, model capacity, and training dynamics affect memorization across various GLM architectures. The study demonstrates that GLMs do memorize training data to varying degrees and that no single attack method fully captures this risk, highlighting the necessity of multi-vector privacy auditing for genomic AI systems.
arXiv:2606. 08945v1 Announce Type: new Abstract: We investigate whether information about time-to-event risk estimated by a Cox proportional hazards model can be transferred into a generative large language model.
The article presents a new semantic model for representing scientific evidence, specifically tailored to genetics, that extends existing standards by adding fine‑grained, domain‑specific structure. It aligns with FHIR Evidence and SEPIO, incorporates a compact vocabulary validated by SHACL, and was tested in a human‑AI annotation pilot on six genetics papers, producing 28 evidence items and 95 source‑anchored assertions. The authors argue that this model advances trustworthy, AI‑ready infrastructure for variant interpretation by providing a reference data model and validation schema for genetic evidence.
arXiv:2608. 00935v1 Announce Type: new Abstract: Electronic Health Records (EHRs) are widely used for clinical risk prediction using machine learning.