arXiv Machine Learning By Alberto Acedo

Omega-N: Interpretable Structural Node Descriptors and Their Applicability Domain

Read the original on arXiv Machine Learning →

The paper introduces Omega‑N, a set of ten interpretable node‑level structural descriptors derived from localizing four factors of a composite structural index. By correcting the ill‑conditioned localization with a configuration‑null excess and a multi‑scale personalized‑PageRank neighbourhood, Omega‑N achieves competitive or superior performance in six in‑domain node‑classification tasks compared to a recursive feature engine that uses up to 252 features. In drug‑target prioritisation on protein interaction networks, Omega‑N improves AUPRC by 0.073 to 0.144 over a centrality baseline and remains robust across independent datasets and bias controls, though it offers no benefit when combined with Node2Vec. whyItMatters":"The study demonstrates that a compact, interpretable set of structural features can match or exceed more complex feature sets in practical graph‑based prediction tasks, particularly in biomedical network analysis."

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 15

Scalable partial information decomposition for symptom networks via supervised embeddings

The paper introduces ePID, an embedding-based approach that scales partial information decomposition (PID) to large symptom networks by compressing non‑focal symptoms into a low‑cardinality discrete embedding. Using a supervised Agglomerative Conditional Information Bottleneck (ACIB) embedding, ePID accurately recovers source‑unique, remainder‑unique, redundant, and synergistic components for each ordered source‑target pair across 83 real‑world datasets, outperforming 12 other embeddings. Applied to PHQ‑9 and the Interpersonal Reactivity Index, ePID reveals distinct patterns of redundancy and synergy that align with each instrument’s construction, demonstrating its ability to separate overlapping from interaction‑dependent information in symptom networks.

By Cillian Hourican, Eric Dignum, Rick Quax, Debraj Roy
arXiv Machine Learning
5d ago

Interpretable-by-Design Descriptor Portfolios Match a 2048-Dimensional Foundation Embedding on Low-Data Molecular Assays

The study evaluates whether a portfolio of compact, semantically named descriptor blocks can match the performance of a 2048‑dimensional CheMeleon embedding in low‑data molecular assays. Using a fixed 11‑dimensional physicochemical base and greedily adding provenance‑screened blocks, the portfolio achieves a mean test AUC of 0.762 across nine ADME/Tox assays, comparable to CheMeleon’s 0.764 and better than Mordred’s 0.756. The results meet a predeclared pooled parity threshold but not all per‑assay thresholds, and further analysis confirms the competitiveness of the auditable representation while highlighting unresolved assay‑level differences.

By Yiqi Yao, Miquel Duran-Frigola
arXiv Machine Learning
Jun 11

GraphInfer-Bench: Benchmarking LLM's Inference Capability on Graphs

arXiv:2606. 11562v1 Announce Type: new Abstract: Graph analysis underlies many applications whose answers cannot be looked up in a single record or retrieved along a path: laundering rings, drug repurposing, user preference, and scientific theme are all inferred from a node together with its neighbourhood.

By Zhuoyi Peng, Jingzhou Jiang, Hanlin Gu, Lixin Fan, Yi Yang