arXiv:2608. 09236v1 Announce Type: new Abstract: Federated learning enables privacy-preserving collaboration across distributed devices without centralizing local data.
By Jaeheon Kim, Hokeun Kim, Bong Jun Choi
arXiv:2608. 07881v1 Announce Type: new Abstract: Clustering mixed tabular data requires a unified metric space to bridge the inherent heterogeneity between continuous numerical measurements and discrete categorical symbols.
By Zihua Yang, Zhencheng Xie, Junyang Chen, Liang Xie, Yiqun Zhang, Mengke Li, Yang Lu
arXiv:2606. 08090v1 Announce Type: cross Abstract: Evaluating a natural-language yes/no predicate over a document corpus under an accuracy target - the semantic filter - is a cornerstone of LLM-based data processing.
By Kyoungmin Kim, Martin Catheland, Anastasia Ailamaki
arXiv:2606. 11499v1 Announce Type: cross Abstract: The performance of modern language models depends critically on pretraining data composition.
By Vedant Badoni, Danqi Chen, Xinyi Wang
arXiv:2605. 08759v3 Announce Type: replace Abstract: Existing granular-ball generation methods are still mainly driven by handcrafted quality measures and heuristic splitting or stopping criteria, which may weaken the transparency of local generation decisions in clustering.
By Zeqiang Xian, Caihui Liu, Yong Zhang, Wenjing Qiu, Duoqian Miao, Witold Pedrycz
arXiv:2405. 07780v3 Announce Type: replace-cross Abstract: This paper explores test-agnostic long-tail recognition, a challenging long-tail task where the test label distributions are unknown and arbitrarily imbalanced.
By Zhiyong Yang, Qianqian Xu, Sicong Li, Zitai Wang, Xiaochun Cao, Qingming Huang