arXiv Machine Learning

Preserving Clusters in Error-Bounded Lossy Compression of Scientific Particle Data

arXiv:2604. 18801v2 Announce Type: replace Abstract: Scientific particle simulations in cosmology, molecular dynamics, and fluid dynamics produce large-scale datasets whose storage, movement, and analysis increasingly rely on lossy compression.

arXiv Machine Learning
Sep 4

Generative Nested Sampling of Atomistic Thermodynamic Landscapes

The paper introduces NS‑Flows, a flow‑based nested sampling method that replaces Markov‑chain updates with a conditional normalizing flow trained on live sets. By applying this technique to a Lennard‑Jones particle system, the authors achieve over two orders of magnitude fewer energy evaluations and a roughly one‑third reduction in wall‑clock time compared to traditional nested sampling. The study also shows that the flow’s generation efficiency varies non‑monotonically along the annealing trajectory, providing a diagnostic of the system’s internal mode complexity and identifying liquid‑like ensembles as the most challenging for current flow architectures.

By Alessandro Coretti, Nico Unglert, Sebastian Falkner, Georg K. H. Madsen, Christoph Dellago
arXiv Machine Learning
Jun 8

ScatterPrism: convergence for generative simulation and inverse problems in particle and nuclear physics

arXiv:2604. 01313v2 Announce Type: replace Abstract: High-fidelity simulations and complex inverse problems, such as detector modeling and unfolding, are computationally intensive bottlenecks across subatomic physics, yet essential for accurate physical interpretation.

By Zeyu Xia, Tyler Kim, Trevor Reed, Judy Fox, Geoffrey Fox, Adam Szczepaniak
arXiv Machine Learning
Sep 17

Bracketing Uncertainty in Clustering Under the Manifold Hypothesis

The paper formalizes a geometric tradeoff between ambient separation and sampling gaps to determine when distinct manifold components can be reliably separated in clustering. It introduces a threshold phenomenon for mutual‑k‑nearest‑neighbor graphs, defining an uncertainty zone where the number of clusters cannot be identified. The authors propose Manifold‑Based Clustering (MBC), which outputs a bracket interval quantifying this uncertainty rather than forcing a single cluster count.

By Savik Kinger, Luciano Dyballa, Steven W. Zucker
arXiv Machine Learning
2d ago

T-ARC: Topology-Aware Randomized Clustering via Distributionally Robust Stochastic Block Models

The paper introduces T-ARC, a clustering algorithm that integrates topological information into the K‑means objective by coupling a data‑fidelity term with a graph‑cut penalty. The latent graph is modeled as a random realization from a Stochastic Block Model, whose parameter is optimized via Distributionally Robust Optimization, using a persistence‑based similarity matrix derived from zero‑dimensional persistent homology. Experiments on synthetic non‑convex data and Fashion‑MNIST subsets demonstrate that T‑ARC recovers latent topological structures and outperforms K‑means on curved and interleaved clusters while remaining competitive and more stable on real data.

By Serena Grazia De Benedictis, Andersen Ang, Nicoletta Del Buono, Flavia Esposito, Laura Selicato