arXiv:2605. 18629v2 Announce Type: replace Abstract: Sparse autoencoders (SAEs) are one of the main methods to interpret the inner workings of deep neural networks (DNNs), decomposing activations into higher-dimensional features.
By Micha{\l} Brzozowski, Neo Christopher Chung
arXiv:2608.07632v2 Announce Type: replace-cross
Abstract: Image-based profiling captures rich phenotypic signatures for drug discovery and functional genomics. Large public datasets like JUMP Cell Pa...
By Al\'an F. Mu\~noz, Johan Fredin Haslum, Runxi Shen, Anne E. Carpenter, Shantanu Singh
MODIS is a semi‑supervised framework for integrating multi‑omics data that are often unpaired, partially labeled, and scarce, such as in rare disease studies. It trains on a large reference database and a small target dataset simultaneously, using diagonal integration and class‑label alignment to handle class imbalance. The architecture combines variational auto‑encoders, a class classifier, and an adversarially trained modality classifier, with a regularized relativistic GAN loss for stable training, and demonstrates high accuracy on synthetic data and the TCGA cancer dataset.
By Daniel Lepe-Soltero, Thierry Arti\`eres, Ana\"is Baudot, Paul Villoutreix
arXiv:2512. 17678v2 Announce Type: replace-cross Abstract: Selecting compact and informative gene subsets from single-cell transcriptomic data is essential for biomarker discovery, improving interpretability, and cost-effective profiling.
By Daphn\'e Chopard, Jorge da Silva Gon\c{c}alves, Irene Cannistraci, Thomas M. Sutter, Julia E. Vogt
arXiv:2608. 06993v1 Announce Type: cross Abstract: Large-scale pretrained time-series models achieve strong results through large-scale pretraining and task-agnostic representation learning, but they rely on abundant, diverse data that industrial and scientific domains often lack.
By Gregor Molan (Comtrade 360 d.o.o., Letali\v{s}ka cesta 29b, Ljubljana, 1000, Slovenia), Grafika Jati (Comtrade 360 d.o.o., Letali\v{s}ka cesta 29b, Ljubljana, 1000, Slovenia), Francesco Barchi (Alma Mater Studiorum - Universita di Bologna, Department of Electrical, Electronic, and Information Engineering), Andrea Acquaviva (Alma Mater Studiorum - Universita di Bologna, Department of Electrical, Electronic, and Information Engineering), Alja\v{z} Osterman (LE-Tehnika d.o.o., \v{S}uceva 27, Kranj, 4000, Slovenia), Martin Molan (Comtrade AI GmbH, Grafenauweg 8, Zug, 6300, Switzerland)
arXiv:2608. 08182v1 Announce Type: cross Abstract: Machine learning models for MALDI-TOF mass spectrometry have shown considerable promise for clinical microbiology tasks such as microbial identification and antimicrobial resistance prediction.
By Alejandro L. Garc\'ia-Navarro, Carlos Sevilla-Salcedo, Bel\'en Rodr\'iguez-S\'anchez, Vanessa G\'omez-Verdejo
CellMSA introduces a novel single‑cell representation learning framework that leverages a multiple‑sequence‑alignment‑inspired context model. For each target cell, it retrieves relevant cells across batches and related cell types, summarizing cross‑cell patterns into a context‑dependent gene‑pair representation that is fed into a pair‑aware encoder. Pretraining on a massive human single‑cell corpus (≈109 million cells) and subsequent benchmarks demonstrate consistent performance gains over existing methods.
By Suyuan Zhao, Minghao Liu, Yizhen Luo, Zaiqing Nie
arXiv:2608. 06727v1 Announce Type: new Abstract: Transformer models for high-dimensional omics analysis process thousands of genes or pathways, although only a subset requires deep computation.
By Koushik Howlader, Tirtho Roy, Md Tauhidul Islam, Wei Le
arXiv:2602. 17330v5 Announce Type: replace-cross Abstract: Comparative analysis of adaptive immune repertoires at population scale is hampered by two practical bottlenecks: the near-quadratic cost of pairwise affinity evaluations and dataset imbalances that obscure clinically important minority clonotypes.
By Rong Fu, Zijian Zhang, Kun Liu, Jiekai Wu, Xianda Li, Simon Fong
arXiv:2608.28923v1 Announce Type: cross
Abstract: Data augmentation is a cornerstone of deep learning pipelines, yet existing strategies treat it as a static, model-agnostic preprocessing step, eithe...
By Noah Videcrantz, Mostafa Mehdipour Ghazi
arXiv:2603. 13377v2 Announce Type: replace-cross Abstract: Representation learning has driven major advances in natural image analysis by enabling models to acquire high-level semantic features.
By Ivan Svatko, Maxime Sanchez, Ihab Bendidi, Gilles Cottrell, Auguste Genovesio
arXiv:2608. 02690v1 Announce Type: new Abstract: On-device training of deep neural networks is fundamentally constrained by the computational and memory costs of large-scale datasets.
By Hetian Liu, Jin Cui, Mengcheng Shi, Yanbin Hu, Xinyue Long, Boran Zhao, Pengju Pen