arXiv:2606. 13754v1 Announce Type: new Abstract: Anomaly detection is a fundamental component of intelligent systems with applications in healthcare, cybersecurity, smart grids, and IoT environments.
By Ghazal Ghajari, Elaheh Ghajari, Ashutosh Ghimire, Saeid Ataei, Faris Alsulami, Fathi Amsaad
arXiv:2407. 00809v4 Announce Type: replace Abstract: This paper introduces the Kernel Neural Operator (KNO), a provably convergent operator-learning architecture that utilizes compositions of deep kernel-based integral operators for function-space approximation of operators (maps from functions to functions).
By Matthew Lowery, John Turnage, Zachary Morrow, John D. Jakeman, Akil Narayan, Shandian Zhe, Varun Shankar
arXiv:2203. 04711v2 Announce Type: replace Abstract: We present a framework for embedding graph structured data into a vector space, taking into account node features and topology of a graph into the optimal transport (OT) problem.
By Dai Hai Nguyen, Koji Tsuda
The paper introduces Backward Kernel Herding, an algorithm that iteratively removes data points to create representative subsets for kernel learning, achieving performance comparable to state‑of‑the‑art methods while speeding up subsampling when the reduced size is less than half the original dataset. It also proposes Flexible Kernel Thinning, an extension that allows construction of subsets of any size, not just successive halvings, and demonstrates that this method often yields the best predictive performance. Experiments on Gaussian Processes and Kernel Support Vector Machines show that Backward Kernel Herding excels in training‑time efficiency, while Flexible Kernel Thinning offers superior predictive accuracy and competitive memory usage, emphasizing the need to choose a reduction strategy based on the desired trade‑off between performance, cost, and memory.
By Blanca Cano-Camarero, Yago R. Aguado-Carrillo-de-Albornoz, \'Angela Fern\'andez-Pascual, Jos\'e R. Dorronsoro
The paper introduces a transition-based derandomization framework for dense binary hypervector codebooks used in hyperdimensional computing. It targets two similarity families—exponential and linear decay with scalar separation—and separates the similarity law, derandomization variant, and generator construction. The authors formalize variants that constrain initial Hamming weight, update-count variability, and update balance, deriving exact finite-dimensional expressions for bias, variance, and RMS error, and validate the theory with simulations to guide practical codebook design.
By Dmitri Rachkovskij, Evgeny Osipov, Olexander Volkov, Denis Kleyko, Vaclav Snasel
arXiv:2606. 13871v1 Announce Type: new Abstract: Tabular data embeddings have become a cornerstone of data profiling and data integration pipelines, enabling tasks such as entity annotation and resolution; schema matching; column type detection; and table search, among others.
By Sebasti\'an Bugedo, Stijn Vansummeren
arXiv:2608. 01528v1 Announce Type: new Abstract: Vector symbolic architectures (VSA) are widely used for reasoning in neuro-symbolic (NeSy) AI, yet high-dimensional codebooks often create severe memory bottlenecks that limit scalability and deployment.
By Weilun Wang, Wantong Li
The paper introduces the Sparse Landmark Embedding (SLE) kernel, a new framework that removes the need for conditionally negative definite (CND) distance measures in kernel methods and Gaussian Processes. By embedding each input into a sparse feature vector using compactly supported bump functions centered at all training points, any standard positive semi-definite (PSD) kernel can be applied in this embedding space, guaranteeing PSD for arbitrary distance measures. The authors provide theoretical guarantees on PSD, sparsity, stability, and universal approximation, and show through experiments with geodesic and Wasserstein distances that the SLE kernel matches or surpasses domain-specific baselines in predictive accuracy and uncertainty quantification.
By Marcus M. Noack, Maher B. Alghalayini, Mark D. Risser
arXiv:2510. 00566v4 Announce Type: replace-cross Abstract: Approximate Nearest-Neighbor Search (ANNS) pipelines for high-dimensional neural embeddings spend the bulk of their query time in candidate verification, making it the primary bottleneck in the search process.
By Vansh Ramani, Alexis Schlomer, Akash Nayar, Sayan Ranu, Jignesh M. Patel, Panagiotis Karras
arXiv:2605. 14981v2 Announce Type: replace Abstract: Gromov--Wasserstein (GW) distances compare graphs, shapes, and point clouds through internal distances, without requiring a common coordinate system.
By Ao Xu, Tieru Wu
arXiv:2603. 05500v2 Announce Type: replace-cross Abstract: Efficient and stable training of large language models (LLMs) remains a core challenge in modern machine learning systems.
By Zeju Qiu, Lixin Liu, Adrian Weller, Han Shi, Weiyang Liu
arXiv:2411. 03253v2 Announce Type: replace-cross Abstract: We propose a general framework for end-to-end learning of data structures.
By Omar Salemohamed, Laurent Charlin, Shivam Garg, Vatsal Sharan, Gregory Valiant