Hugging Face Trending Papers

New Benchmarking Shows Limited Generalization Power of TCR Antigenic Epitope Prediction Models

Accurate computational prediction of T cell receptor (TCR) antigen specificity would transform the study of T cell biology and enable scalable immune engineering, yet existing models lack sufficient sensitivity and specificity for broad applications. A major limitation is the absence of rigorously defined, unseen benchmark datasets that allow unbiased evaluation of model performance and generalizability.

arXiv Machine Learning
Jun 4

New Benchmarking Shows Limited Generalization Power of TCR Antigenic Epitope Prediction Models

arXiv:2606. 04994v1 Announce Type: new Abstract: Accurate computational prediction of T cell receptor (TCR) antigen specificity would transform the study of T cell biology and enable scalable immune engineering, yet existing models lack sufficient sensitivity and specificity for broad applications.

By Yiming Liao, Yiheng Li, Ning Jiang, Bo Li, Keke Chen
arXiv AI
Sep 25

CaliPPer: quantifying, predicting and improving AI model performance for binding prediction

CaliPPer is a post‑hoc framework that calibrates and predicts the performance of binding‑prediction models by combining a multi‑chain Sample‑to‑Domain Distance (S2DD) metric with distance‑aware Bayesian recalibration. It operates at three resolutions—generalisability score, aggregate performance prediction, and per‑sample confidence—achieving strong distance‑performance correlations (|r| = 0.80–0.92) and low prediction errors for AUROC, AP, and F1. In retrospective analyses of five published studies, CaliPPer increased true discovery rates, improving AUROC by up to +0.20 on unseen epitopes and variants and raising confirmed neoantigen findings from 0/5 to 3/5.

By Jian-Qing Zheng, Hantao Lou, Zinan Yin, Sam Farrar, Yuze Zhou, Elie Antoun, Xiangxi Wang, Xuetao Cao, Tao Dong
arXiv AI
4d ago

Explainability from Training with Applications to TCR-Epitope Prediction

The paper introduces Explainability from Training (EFT), a model‑agnostic method that tracks how deep learning models learn and organize evidence during training. EFT is applied to four leading T cell receptor‑epitope prediction models, revealing distinct learning trajectories for CNNs and transformers, conflicts between TCR alpha and beta chain evidence, and differences in feature preferences when using real versus predicted structural data. The authors also present a new benchmark, TCR‑XAI2, comprising 388 experimentally resolved TCR‑epitope structures and several predicted models to evaluate these insights.

By Jiarui Li, Zixiang Yin, Samuel Landry, Zhengming Ding, Ramgopal Mettu
arXiv Machine Learning
Jun 30

Transformer-Based Active Learning for Data-Efficient Vaccine Epitope Selection in PRRS

arXiv:2606. 28659v1 Announce Type: cross Abstract: High-fidelity molecular docking simulations can produce biologically relevant estimates of epitope-receptor binding affinity but are computationally expensive and therefore limit the number of candidates that can be screened for vaccine design.

By Aspen Erlandsson Brisebois, Zahed Khatooni, Connor Burbridge, Brook Byrns, Heather L. Wilson, Sureesh Tikoo, Steven Rayan, Gordon Broderick
arXiv Machine Learning
Jul 23

SwiftRepertoire: Few-Shot Immune-Signature Synthesis via Dynamic Kernel Codes

arXiv:2602. 01051v5 Announce Type: replace Abstract: Repertoire-level analysis of T cell receptors offers a biologically grounded signal for disease detection and immune monitoring, yet practical deployment is impeded by label sparsity, cohort heterogeneity, and the computational burden of adapting large encoders to new tasks.

By Rong Fu, Muge Qi, Yang Li, Yabin Jin, Jiekai Wu, Chunlei Meng, Juntao Gao, Li Bao, Qi Zhao, Wei Luo, Youjin Wang, Simon Fong
arXiv Machine Learning
Sep 15

An immune world model for multiscale forecasting and therapeutic hypothesis generation

arXiv:2609.14709v1 Announce Type: new Abstract: Immune therapies act across cell-intrinsic programs, tissue ecosystems, and patient-specific immune states, yet most predictors address these scales se...

By Taoyong Cui, Xi Wang, Zonghang Li, Jinchao Ding, Lingsen You, Yuzhi Xu, Wanghan Xu, Fang Wu, Kejun Ying, Wanli Ouyang, Pheng Ann Heng, Ling Yang, Zhenfei Yin, Yingcheng Wu
arXiv Machine Learning
4d ago

Estimating the Causal Effects of T Cell Receptors

The paper introduces a method for estimating the causal effects of T cell receptor (TCR) sequences on patient outcomes using observational TCR sequencing and clinical data. It corrects for unobserved confounders by leveraging the pre-selection TCR repertoire generated through V(D)J recombination as a natural experiment, and employs permutation‑invariant neural networks to scale to millions of sequences. The approach is validated on semisynthetic data and applied to COVID‑19 severity, identifying TCRs that are observed in patients, bind SARS‑CoV‑2 antigens in vitro, and positively influence clinical outcomes.

By Eli N. Weinstein, Elizabeth B. Wood, David M. Blei
arXiv Machine Learning
Sep 18

Transcriptomic Models for Immunotherapy Response Prediction Show Limited Cross-cohort Generalisability

The study evaluated nine transcriptomic models—five bulk RNA‑seq and four single‑cell RNA‑seq—designed to predict response to immune checkpoint inhibitors. Across independent datasets, bulk models performed near chance while single‑cell models offered only modest gains, and pathway analyses revealed inconsistent biomarker signals. The results highlight the limited cross‑cohort robustness and biological consistency of current transcriptomic ICI predictors.

By Yuheng Liang, Lucy Chhuo, Ahmadreza Argha, Nona Farbehi, Lu Chen, Roohallah Alizadehsani, Mehdi Hosseinzadeh, Min Yang, Thantrira Porntaveetusm, Youqiong Ye, Hamid Alinejad-Rokny
arXiv AI
Jul 23

SubQuad: Near-Quadratic-Free Structure Inference with Distribution-Balanced Objectives in Adaptive Receptor framework

arXiv:2602. 17330v5 Announce Type: replace-cross Abstract: Comparative analysis of adaptive immune repertoires at population scale is hampered by two practical bottlenecks: the near-quadratic cost of pairwise affinity evaluations and dataset imbalances that obscure clinically important minority clonotypes.

By Rong Fu, Zijian Zhang, Kun Liu, Jiekai Wu, Xianda Li, Simon Fong