arXiv:2609.37604v1 Announce Type: new
Abstract: Graph foundation models need a discrete token representation, but casting a graph as a generatable token sequence faces a structural obstacle: edges sp...
By Yuxiang Yao, Zijun Zhao
The paper introduces ePID, an embedding-based approach that scales partial information decomposition (PID) to large symptom networks by compressing non‑focal symptoms into a low‑cardinality discrete embedding. Using a supervised Agglomerative Conditional Information Bottleneck (ACIB) embedding, ePID accurately recovers source‑unique, remainder‑unique, redundant, and synergistic components for each ordered source‑target pair across 83 real‑world datasets, outperforming 12 other embeddings. Applied to PHQ‑9 and the Interpersonal Reactivity Index, ePID reveals distinct patterns of redundancy and synergy that align with each instrument’s construction, demonstrating its ability to separate overlapping from interaction‑dependent information in symptom networks.
By Cillian Hourican, Eric Dignum, Rick Quax, Debraj Roy
The study evaluates whether a portfolio of compact, semantically named descriptor blocks can match the performance of a 2048‑dimensional CheMeleon embedding in low‑data molecular assays. Using a fixed 11‑dimensional physicochemical base and greedily adding provenance‑screened blocks, the portfolio achieves a mean test AUC of 0.762 across nine ADME/Tox assays, comparable to CheMeleon’s 0.764 and better than Mordred’s 0.756. The results meet a predeclared pooled parity threshold but not all per‑assay thresholds, and further analysis confirms the competitiveness of the auditable representation while highlighting unresolved assay‑level differences.
By Yiqi Yao, Miquel Duran-Frigola
arXiv:2603. 24025v2 Announce Type: replace Abstract: Unsupervised learning of high-dimensional data is challenging due to irrelevant or noisy features obscuring underlying structures.
By Chen Ma, Wanjie Wang, Shuhao Fan
arXiv:2407. 07357v3 Announce Type: replace Abstract: Predicting signed interactions in biological networks is crucial for understanding drug mechanisms and facilitating drug repurposing.
By Ziye Zhou, Meijie Wang, Lun Yu
arXiv:2606. 11562v1 Announce Type: new Abstract: Graph analysis underlies many applications whose answers cannot be looked up in a single record or retrieved along a path: laundering rings, drug repurposing, user preference, and scientific theme are all inferred from a node together with its neighbourhood.
By Zhuoyi Peng, Jingzhou Jiang, Hanlin Gu, Lixin Fan, Yi Yang
The paper investigates a failure mode in Graph-JEPA, a joint‑embedding predictive model trained on a large scientific‑reasoning graph. Despite achieving high linear‑probe accuracy and effective rank, the learned representation contains almost no usable instance information, as shown by retrieval metrics. The authors diagnose the issue to variance allocation in the objective, propose a repair that restores near‑perfect information recovery, and demonstrate that the problem persists even after repair, highlighting limitations in the evaluation metrics used.
By Gollam Rabby, S\"oren Auer
arXiv:2606. 17579v1 Announce Type: cross Abstract: Adding LLM-generated node features to graph neural networks (GNNs) is widely reported to improve accuracy on standard benchmarks.
By Zhongyuan Wang, Pratyusha Vemuri
The paper introduces Murmur2Vec, a lightweight, alignment‑free embedding that uses k‑mer counts hashed with MurmurHash to create a compact representation for biological sequences. It provides a full theoretical analysis, including bias/variance formulas, a Johnson–Lindenstrauss‑style concentration bound, and an excess‑risk bound that clarifies the trade‑off between hash‑table size and classifier performance. Empirically, Murmur2Vec matches or surpasses a fine‑tuned 650M‑parameter ESM‑2 protein language model across several classification tasks, including SARS‑CoV‑2 spike lineage and HIV‑1 Env subtype identification.
By Sarwan Ali, Taslim Murad, Imdadullah Khan, Safi Faizullah
arXiv:2607. 19618v1 Announce Type: cross Abstract: Genomic language models achieve strong performance across regulatory-genomics tasks, yet what these models internally represent remains opaque, and the field lacks a principled procedure for verifying that an apparent ``concept'' inside a model is real rather than an artifact of sequence composition.
By Sarwan Ali
arXiv:2605. 15511v2 Announce Type: replace Abstract: Graph Neural Networks (GNNs) have become the dominant framework for inductive graph-level learning.
By Louisa Cornelis, Johan Mathe, Louis Van Langendonck, Guillermo Bern\'ardez, Nina Miolane
arXiv:2606. 03002v2 Announce Type: replace-cross Abstract: Quantization is a standard path to deploying large language models, and quantized models are typically judged acceptable when perplexity or downstream accuracy remains close to the full-precision original.
By Evan Duan