Cluster Analysis with Resampling for Validation and Exploration (CARVE)
arXiv:2606. 00327v1 Announce Type: cross Abstract: Clustering is widely used across the sciences as the foundation for downstream data-driven scientific discoveries.
arXiv:2604. 18801v2 Announce Type: replace Abstract: Scientific particle simulations in cosmology, molecular dynamics, and fluid dynamics produce large-scale datasets whose storage, movement, and analysis increasingly rely on lossy compression.
arXiv:2606. 00327v1 Announce Type: cross Abstract: Clustering is widely used across the sciences as the foundation for downstream data-driven scientific discoveries.
The paper introduces NS‑Flows, a flow‑based nested sampling method that replaces Markov‑chain updates with a conditional normalizing flow trained on live sets. By applying this technique to a Lennard‑Jones particle system, the authors achieve over two orders of magnitude fewer energy evaluations and a roughly one‑third reduction in wall‑clock time compared to traditional nested sampling. The study also shows that the flow’s generation efficiency varies non‑monotonically along the annealing trajectory, providing a diagnostic of the system’s internal mode complexity and identifying liquid‑like ensembles as the most challenging for current flow architectures.
arXiv:2601. 11626v2 Announce Type: replace-cross Abstract: Large collections of matrices arise throughout modern machine learning, signal processing, and scientific computing, where they are commonly compressed by concatenation followed by truncated singular value decomposition (SVD).
arXiv:2609. 30477v1 Announce Type: cross Abstract: Exact Euclidean \(K\)-means partitions \(n\) observations into \(K\) unlabelled clusters, but the unrestricted search is generally exponential.
arXiv:2604. 01313v2 Announce Type: replace Abstract: High-fidelity simulations and complex inverse problems, such as detector modeling and unfolding, are computationally intensive bottlenecks across subatomic physics, yet essential for accurate physical interpretation.
arXiv:2512. 16558v3 Announce Type: replace Abstract: Clustering is a cornerstone of modern data analysis.
arXiv:2608.22065v1 Announce Type: cross Abstract: Whether the slow solar wind originates from one coronal source or two distinct channels remains a central open question in heliophysics. Resolving th...
arXiv:2607. 05469v1 Announce Type: cross Abstract: Unsupervised graph clustering is a fundamental technique for uncovering underlying semantic patterns in large-scale networks.
arXiv:2608.30262v1 Announce Type: new Abstract: Scientific simulations generate collections of physical fields with heterogeneous statistics and dependencies, yet learned compressors often encode tho...
The paper formalizes a geometric tradeoff between ambient separation and sampling gaps to determine when distinct manifold components can be reliably separated in clustering. It introduces a threshold phenomenon for mutual‑k‑nearest‑neighbor graphs, defining an uncertainty zone where the number of clusters cannot be identified. The authors propose Manifold‑Based Clustering (MBC), which outputs a bracket interval quantifying this uncertainty rather than forcing a single cluster count.
arXiv:2609.36074v1 Announce Type: new Abstract: Memory-efficient scaling on clustering problems without sacrificing statistical accuracy is of central interest for large-scale data analysis and machi...
The paper introduces T-ARC, a clustering algorithm that integrates topological information into the K‑means objective by coupling a data‑fidelity term with a graph‑cut penalty. The latent graph is modeled as a random realization from a Stochastic Block Model, whose parameter is optimized via Distributionally Robust Optimization, using a persistence‑based similarity matrix derived from zero‑dimensional persistent homology. Experiments on synthetic non‑convex data and Fashion‑MNIST subsets demonstrate that T‑ARC recovers latent topological structures and outperforms K‑means on curved and interleaved clusters while remaining competitive and more stable on real data.