arXiv:2606. 18209v1 Announce Type: new Abstract: Dataset distillation (DD) has emerged as a prominent approach in data centric machine learning, aiming to synthesize compact training sets for efficient training by compressing the information in large datasets into a small number of synthetic samples.
By Trisha Mittal, Akshay Mehra, Joshua Kimball
arXiv:2509. 05663v4 Announce Type: replace Abstract: Truly unsupervised approaches for time series anomaly detection are rare in the literature.
By Lucas Correia, Jan-Christoph Goos, Thomas B\"ack, Anna V. Kononova
arXiv:2607. 03188v1 Announce Type: cross Abstract: Episode mining aims to extract subsequences of events that possess certain distinctive properties and constitute facts valuable to the user.
By Maxim Ivanov, Matvei Smirnov, Alisa Strazdina, George Chernishev
arXiv:2609.36686v1 Announce Type: new
Abstract: Identifying the root cause of an anomaly among hundreds of sensors is critical for preventing safety incidents and costly downtime in complex monitored...
By Hada Melino Muhammad, Luan Pham, Laure Barri\`ere, Sachin Shetty, Leonardo Pulga, Flora D. Salim
The paper investigates how AI‑generated data, treated as anomalies linked to main data points, influences batch decompositions and undersampling in random datasets. Using redundancy graphs and iterative methods, it derives bounds for the minimum size of a strongly dissimilar decomposition and shows a phase transition: when anomalies are few, the minimum size depends on main data, but beyond a threshold it is dominated by anomalies. Additionally, the authors present a size criticality result for strong similarity in randomly undersampled datasets, illustrated with categorical data examples where the overall space far exceeds the dataset size.
By Ghurumuruhan Ganesan
arXiv:2511. 03217v2 Announce Type: replace-cross Abstract: Large language models (LLMs) excel in generating fluent utterances but can lack reliable grounding in verified information.
By Shaghayegh Kolli, Richard Rosenbaum, Timo Cavelius, Lasse Strothe, Andrii Lata, Jana Diesner
arXiv:2411. 03253v2 Announce Type: replace-cross Abstract: We propose a general framework for end-to-end learning of data structures.
By Omar Salemohamed, Laurent Charlin, Shivam Garg, Vatsal Sharan, Gregory Valiant
arXiv:2607. 04949v1 Announce Type: new Abstract: We study the problem of k-means clustering on large datasets.
By Cristian Boldrin, Fabio Vandin
arXiv:2607. 25135v1 Announce Type: new Abstract: Recent advances in RAG aim to optimize for performance by paying high ingestion costs for knowledge ingestion: building knowledge graphs or extracting SQL tables.
By Hilaf Hasson, Aditya Chakravarty, Jayant Thomas, Krishna Gogineni
arXiv:2608. 14999v1 Announce Type: cross Abstract: Constructing special graphs is an important task within graph theory and computer science.
By Zohair Raza Hassan, Deepak Pandita
arXiv:2607. 01647v1 Announce Type: cross Abstract: Data science aims to derive actionable insights from heterogeneous raw data, unlocking the value of the massive amounts of data generated in modern society.
By Zhaoyan Sun, Shan Zhong, Daizhou Wen, Jiaxing Han, Guoliang Li, Ying Yan, Peng Zhang, Yu Su, Xiang Qi, Baolin Sun, Chengyuan Yang, Tao Fang, Huaiyu Ruan
arXiv:2607. 20163v1 Announce Type: cross Abstract: The rapid growth of biomedical knowledge has made the validation of automatically generated biological annotations a major bottleneck in biomedical curation.
By Emanuele Cavalleri, Miad Alavinezhad, Dario Malchiodi, Marco Mesiti