EXPOSE is a framework that applies Sparse Autoencoders to Vision Foundation Model embeddings in computational pathology, aiming to separate biological signals from domain‑specific noise. By training a sparse representation of VFM features and using a linear classifier to flag domain‑specific latent dimensions, the method masks these components before downstream relapse prediction, avoiding the need to retrain the backbone model. Experiments on a large prostate cancer dataset demonstrate that removing domain‑specific features improves cross‑domain performance and raises the Domain Robustness Index (DoRI).
By Anja Witte, Maximilian Lennartz, Jan Baumbach, Guido Sauter, Stefan Bonn, Patrick Fuhlert, Marina Zimmermann
arXiv:2606. 06233v1 Announce Type: cross Abstract: Principal component analysis (PCA) is one of the most widely used unsupervised dimension reduction techniques.
By Benedikt Seiter, Anya Fries, Julius von K\"ugelgen, Jonas Peters
arXiv:2507.23559v2 Announce Type: replace-cross
Abstract: Certain data are naturally modeled by networks or weighted graphs, be they biological networks or mobility networks. When there is no canonic...
By Elodie Maignant, Xavier Pennec, Alain Trouv\'e, Anna Calissano
arXiv:2606. 03553v1 Announce Type: cross Abstract: While principal component analysis (PCA) is a fundamental tool for dimensionality reduction, its dense representations make it ill-suited for high-dimensional data.
By David V\"avinggren, Francis Bach, Andr\'e M. H. Teixeira, Dave Zachariah, Ant\^onio H. Ribeiro
arXiv:2605. 20689v2 Announce Type: replace-cross Abstract: High-dimensional language-model embeddings increase storage and search costs, while supervised compressors can overfit when relevance labels are scarce.
By Dongfang Zhao
arXiv:2609.37659v1 Announce Type: cross
Abstract: There has been significant work on understanding the In-Context Learning capabilities of Large Language Models, especially on the induction circuit....
By Adhemar de Senneville, Xavier Bou, J\'er\'emy Anger, Rafael Grompone, Gabriele Facciolo
There has been significant work on understanding the In-Context Learning capabilities of Large Language Models, especially on the induction circuit. For a few-shot classification task, the induction c...
arXiv:2605. 29987v2 Announce Type: replace Abstract: Although multi-scales representation learning enables elastic-dimension embeddings, nested subspaces often suffer from dimensional redundancy and spectral collapse.
By Dang Nguyen Hong, Nhi Ngoc-Yen Nguyen, Huy-Hieu Pham
Introduction Heavy computation is a well-known problem in various ML algorithms today, especially when generative AI is applied to text, images, and other unstructured data. One of the principal approaches to mitigate this problem is to compress input data into a lower-dimensional representation while preserving the main context.
By Vyacheslav Efimov
arXiv:2407. 21311v2 Announce Type: replace-cross Abstract: Unsupervised domain adaptation (UDA) aims to mitigate domain shift, where the distribution of labeled source data differs from that of unlabeled target data.
By Ali Abedi, Q. M. Jonathan Wu, Ning Zhang, Farhad Pourpanah
arXiv:2503.10685v3 Announce Type: replace
Abstract: Unsupervised Domain Adaptation (UDA) enables strong generalization from a labeled source domain to an unlabeled target domain, often with limited d...
By Brun\'o B. Englert, Gijs Dubbelman
arXiv:2601. 19179v2 Announce Type: replace Abstract: Autoencoders have long been considered a nonlinear extension of Principal Component Analysis (PCA).
By Qipeng Zhan, Zhuoping Zhou, Zexuan Wang, Li Shen