arXiv:2509. 03373v2 Announce Type: replace Abstract: Dimensionality reduction methods such as t-SNE and UMAP are popular methods for visualizing data with a potential (latent) clustered structure.
By Elizabeth Coda, Ery Arias-Castro, Gal Mishne
arXiv:2605. 23540v2 Announce Type: replace Abstract: Dimensionality Reduction (DR) methods are widely used to visualize high-dimensional data.
By Diede P. M. van der Hoorn, Alessio Arleo, Fernando V. Paulovich
The paper introduces an unsupervised framework that merges manifold learning with rank‑based interpretable graph embeddings to address the Geometric and Interpretability Gaps in visual representation learning. By first analyzing contextual information on the dataset manifold and then producing sparse, self‑explainable embeddings, the method achieves dimensionality reduction while preserving or improving performance in image retrieval and semi‑supervised Graph Convolutional Network classification. Experiments across varied datasets confirm that these context‑aware representations maintain high downstream effectiveness.
By Thiago C\'esar Castilho Almeida, Gustavo Rosseto Let\'icio, Vinicius Atsushi Sato Kawai, Daniel Carlos Guimar\~aes Pedronette
arXiv:2607. 03978v1 Announce Type: cross Abstract: Low-dimensional projections support interactive visual analysis of high-dimensional data embeddings, but their structure often does not align with analyst-defined semantic relationships.
By Wei Liu, Eric Krokos, Kirsten Whitley, Rebecca Faust, Chris North
arXiv:2606. 16379v1 Announce Type: new Abstract: Evaluating representation similarity is fundamental to representation learning.
By Diogo Soares, Pankhil Gawade, Andrea Dittadi, Ewa Szczurek
arXiv:2608.29001v1 Announce Type: new
Abstract: In a data-driven world, efficiently organizing and mapping relationships between objects is crucial. Graphs are powerful tools for modeling these conne...
By Thiago C\'esar Castilho Almeida, Gustavo Rosseto Let\'icio, Lucas Pascotti Valem, Andr\'e Freitas, Daniel Carlos Guimar\~aes Pedronette
arXiv:2607. 15018v1 Announce Type: cross Abstract: High-dimensional categorical data arise in genetics, biomedicine, and the social sciences, yet visualization tools for such data remain far less developed than those for continuous variables.
By Chun-houh Chen, Shun-Chuan Chang, Chiun-How Kao, Yi-Ju Lee, Shang-Ying Shiu, Yin-Jing Tien, ShengLi Tzeng, Han-Ming Wu
arXiv:2608. 11269v1 Announce Type: cross Abstract: Omics datasets, particularly single-cell RNA sequencing data, are high-dimensional, sparse, noisy, and dominated by zero values, making faithful low-dimensional representation challenging.
By Fenosoa Randrianjatovo, Maya Saleh, Simon Girard, Amadou Barry
arXiv:2606. 02172v1 Announce Type: new Abstract: Learning discriminative visual representations from distributed, heterogeneous data is a fundamental challenge in Federated Learning (FL).
By Mario Casado-Diez, Alejandro Dopico-Castro, Ver\'onica Bol\'on-Canedo, Bertha Guijarro-Berdi\~nas
arXiv:2606. 04451v1 Announce Type: new Abstract: Neighbor embedding algorithms reveal correlations in high-dimensional data by constructing an equivalent graph representation in a lower-dimensional space.
By Mohammad Tariqul Islam, Jason W. Fleischer
The paper introduces Mapping the Concept Landscape (MCL), a framework that replaces high‑dimensional feature embeddings with explicit sample‑level graphs of entities, events, and attributes for image‑caption pairs. By aggregating these graphs into a dataset‑level graph, MCL captures the global distribution of semantic concepts and identifies rare concepts. A greedy algorithm then selects samples to maximize coverage of under‑represented concepts, achieving better pruning efficiency and providing a transparent audit trail.
By Dongyue Wu, Tao Ma
arXiv:2606. 31119v1 Announce Type: new Abstract: Graphs are commonly visualized in 2D, where humans readily interpret spatial relationships, yet such layouts often distort higher-dimensional structure.
By Ya Ji (Khoury College of Computer Sciences, Northeastern University, Seattle), Xuefeng Li (Khoury College of Computer Sciences, Northeastern University, Seattle), Timo Brand (School of Computation, Information and Technology, Technical University of Munich, Heilbronn, Germany), Jacob Miller (School of Computation, Information and Technology, Technical University of Munich, Heilbronn, Germany), Peng Zhang (Khoury College of Computer Sciences, Northeastern University, Seattle), Stephen Kobourov (School of Computation, Information and Technology, Technical University of Munich, Heilbronn, Germany), Yifan Hu (Khoury College of Computer Sciences, Northeastern University, Seattle)