ALICE is a foundation model that estimates mutual information (MI) without per‑distribution training. Trained only on synthetic distributions, it acts as an in‑context estimator of rectified‑flow velocity fields, producing MI via a fixed identity that integrates squared differences between joint and conditional fields. The authors validate ALICE on a challenging benchmark and demonstrate its applicability to unseen data in biology, genetics, and neuroscience, achieving performance comparable to neural estimators trained separately for each distribution.
By Giulio Franzese, Simone Rossi, Pietro Michiardi
arXiv:2511. 18945v4 Announce Type: replace Abstract: We propose a fully data-driven approach to designing mutual information (MI) estimators.
By German Gritsai, Megan Richards, Maxime M\'eloux, Kyunghyun Cho, Maxime Peyrard
arXiv:2602. 00329v4 Announce Type: replace-cross Abstract: Reliable data attribution is essential for mitigating bias and reducing computational waste in modern machine learning, with the Shapley value serving as the theoretical gold standard.
By Meng Ding, Zeqing Zhang, Di Wang, Lijie Hu
FEAT is a foundation model designed for extremely large structured data that replaces quadratic self‑attention with a linear‑complexity dual‑axis encoding architecture. It combines an adaptive‑fusion bidirectional state‑space model with convolutional gated linear attention to achieve permutation‑invariant representation learning in O(N) time. Experiments on 12 real‑world database benchmarks show that FEAT outperforms existing structured data foundation models on zero‑shot tasks and can be up to 50× faster in inference latency.
By Zhenghang Song, Tang Qian, Lu Chen, Yushuai Li, Zhengke Hu, Bingbing Fang, Yumeng Song, Junbo Zhao, Sheng Zhang, Tianyi Li
arXiv:2606. 10798v1 Announce Type: new Abstract: Pretrained time series foundation models (TSFMs) have enabled zero-shot forecasting on unseen target series.
By Yosuke Yamaguchi, Issei Suemitsu, Yuki Kajihara, Wenpeng Wei
GENSCRIPT is an inference‑only pipeline that generates synthetic data without training a generative model. It creates a deterministic statistical profile of the source data, uses a language model to infer field semantics and cross‑column constraints, and then compiles these into an executable sampler that works for single‑table, temporal, and relational data. The method builds generators in minutes, samples large datasets quickly, and achieves fidelity comparable to leading methods while preserving key data relationships such as 1‑to‑1 mappings and primary‑foreign key constraints.
By Zilong Zhao, Abdul Raheem, Jiayu Li, Sohei Arisaka, Darius Lim Hong Yi, Milad Abdollahzadeh, Uzair Javaid, Biplab Sikdar
The paper demonstrates that a tabular foundation model can achieve strong generalization using only a single real table for self‑supervised pre‑training, challenging the belief that large synthetic or real datasets are necessary. By systematically pre‑training and evaluating across diverse benchmarks, the authors show that the number and quality of tasks that can be derived from a dataset are critical for downstream performance. This finding suggests that carefully constructed task sets from limited data can enable effective transfer learning in tabular models.
By Junwei Ma, Nour Shaheen, Alex Labach, Amine Mhedhbi, Frank Hutter, Anthony L. Caterini, Valentin Thomas
arXiv:2605. 23268v2 Announce Type: replace-cross Abstract: In many prediction problems, we have extra information during training (for example, measurements that are expensive or slow to collect) that will not be available when the model is deployed.
By Jiahao Shi, Omar Hagrass, Jason M. Klusowski
arXiv:2509. 14230v2 Announce Type: replace Abstract: While structured pruning presents a highly effective pathway for accelerating Large Language Model (LLM) inference, existing methods frequently suffer from significant performance degradation and demand computationally retraining to recover capabilities.
By Mengting Ai, Tianxin Wei, Sirui Chen, Jingrui He
arXiv:2603.26164v2 Announce Type: replace-cross
Abstract: Data-centric training has emerged as a promising direction for improving large language models (LLMs) by optimizing not only model parameters...
By Hao Liang, Zhengyang Zhao, Mingrui Chen, Meiyi Qiang, Lu Ma, Rongyi Yu, Hengyi Feng, Shixuan Sun, Zimo Meng, Xiaochen Ma, Xuanlin Yang, Qifeng Cai, Ruichuan An, Bohan Zeng, Zhen Hao Wong, Chengyu Shen, Runming He, Zhaoyang Han, Yaowei Zheng, Fangcheng Fu, Conghui He, Bin Cui, Zhiyu Li, Weinan E, Wentao Zhang
arXiv:2509. 20345v3 Announce Type: replace-cross Abstract: The rapid proliferation of high-quality synthetic data -- generated by advanced AI models or collected as auxiliary data from related tasks -- presents both opportunities and challenges for statistical inference.
By Meshi Bashari, Yonghoon Lee, Roy Maor Lotan, Edgar Dobriban, Yaniv Romano
arXiv:2608. 09690v1 Announce Type: new Abstract: Recurrent neural networks (RNNs) are widely used for sequence learning, yet their application is commonly associated with temporal data, although recurrent computation fundamentally operates on ordered sequences rather than on time itself.
By Vagan Terziyan, Artur Terziian, Oleksandra Vitko