arXiv AI By Zhengyang Hu, Yanzhi Chen, Hanxiang Ren, Qunsong Zeng, Youyi Zheng, Adrian Weller, Kaibin Huang, Yanchao Yang

InfoAtlas: A Foundation Model for Zero-Shot Statistical Dependence Estimate

Read the original on arXiv AI →

arXiv:2606. 00241v1 Announce Type: cross Abstract: Measuring statistical dependency between high-dimensional random variables is a fundamental task in data science and machine learning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
4d ago

ALICE: In-context, Zero-shot, Mutual Information Estimation

ALICE is a foundation model that estimates mutual information (MI) without per‑distribution training. Trained only on synthetic distributions, it acts as an in‑context estimator of rectified‑flow velocity fields, producing MI via a fixed identity that integrates squared differences between joint and conditional fields. The authors validate ALICE on a challenging benchmark and demonstrate its applicability to unseen data in biology, genetics, and neuroscience, achieving performance comparable to neural estimators trained separately for each distribution.

By Giulio Franzese, Simone Rossi, Pietro Michiardi
arXiv AI
Jul 23

In-Run Data Shapley for Adam Optimizer

arXiv:2602. 00329v4 Announce Type: replace-cross Abstract: Reliable data attribution is essential for mitigating bias and reducing computational waste in modern machine learning, with the Shapley value serving as the theoretical gold standard.

By Meng Ding, Zeqing Zhang, Di Wang, Lijie Hu
arXiv Machine Learning
Sep 14

FEAT: A Linear-Complexity Foundation Model for Extremely Large Structured Data

FEAT is a foundation model designed for extremely large structured data that replaces quadratic self‑attention with a linear‑complexity dual‑axis encoding architecture. It combines an adaptive‑fusion bidirectional state‑space model with convolutional gated linear attention to achieve permutation‑invariant representation learning in O(N) time. Experiments on 12 real‑world database benchmarks show that FEAT outperforms existing structured data foundation models on zero‑shot tasks and can be up to 50× faster in inference latency.

By Zhenghang Song, Tang Qian, Lu Chen, Yushuai Li, Zhengke Hu, Bingbing Fang, Yumeng Song, Junbo Zhao, Sheng Zhang, Tianyi Li
arXiv Machine Learning
2d ago

Synthesis Without Training: An Inference-Only Pipeline for Tabular, Temporal, and Relational Synthetic Data

GENSCRIPT is an inference‑only pipeline that generates synthetic data without training a generative model. It creates a deterministic statistical profile of the source data, uses a language model to infer field semantics and cross‑column constraints, and then compiles these into an executable sampler that works for single‑table, temporal, and relational data. The method builds generators in minutes, samples large datasets quickly, and achieves fidelity comparable to leading methods while preserving key data relationships such as 1‑to‑1 mappings and primary‑foreign key constraints.

By Zilong Zhao, Abdul Raheem, Jiayu Li, Sohei Arisaka, Darius Lim Hong Yi, Milad Abdollahzadeh, Uzair Javaid, Biplab Sikdar