arXiv Machine Learning

The Economics of Model Collapse: Equilibrium, Welfare, and Optimal Provenance Subsidies in Synthetic Data Markets

arXiv:2605. 20279v2 Announce Type: replace-cross Abstract: Generative artificial intelligence is rapidly transforming the supply side of training data: an increasing share of new tokens, images, and structured records is produced by previous-generation models rather than by human originators.

arXiv Machine Learning
Jul 9

Best-Arm Identification with Generative Proxy

arXiv:2607. 06879v1 Announce Type: new Abstract: Best-arm identification is a canonical model for data-driven decision-making, but in many applications each reward observation is costly.

By Tianyi Ma, Hanzhang Qin, Ruihao Zhu, Jierui Zuo
Hugging Face Trending Papers
Jul 13

The Equilibrium Is the Initialization: Lazy Identity Collapse in Physics-Structured Deep Equilibrium Reasoning

Deep equilibrium models promise input-adaptive implicit computation: harder problems should demand more solver iterations, and the solved equilibrium should encode the result of genuine iterative inference. We report a cautionary study of a port-Hamiltonian DEQ with a learned initialization on two reasoning tasks -- ProofWriter entailment over frozen DeBERTa embeddings and a BFS-verified graph-reachability benchmark -- in which the implicit computation is a silent no-op.

arXiv Machine Learning
Sep 17

Preventing Model Collapse: A Fisher-Rao Perspective on the Dynamics of Training with Synthetic Data

The paper investigates how to avoid model collapse when training large language models with synthetic data. It establishes theoretical guarantees for the minimum ratio of human to synthetic data needed to maintain training stability, using the Fisher‑Rao metric to analyze dynamics on the probability simplex. The authors derive contraction and invariance bounds that remain meaningful even in high dimensions, showing that the required data ratio differs from earlier estimates.

By Matteo Marchi, Jo\~ao Pedro Silvestre, Bahman Gharesifard, Paulo Tabuada
arXiv AI
Aug 26

A tale of perfect fit and phantom optima: how data-driven models can fail in real-time optimization

The paper examines the reliability of data‑driven models for real‑time optimization (RTO) using a vinyl acetate monomer benchmark. Two models—a structured hybrid model and a fully data‑driven neural ODE—accurately reproduce plant measurements but yield economic optima that differ markedly from the plant’s true optimum, producing multiple phantom optima. The study shows that even with noise‑free data and correct initialization, stochastic gradient training can drift to weights that degrade RTO performance, indicating that predictive accuracy alone does not ensure reliable economic outcomes.

By Prithvi Dake, Rahul Bindlish, James B. Rawlings
arXiv AI
Sep 15

Certifiably Interpretable Training of ReLU-MLPs for Boolean Tasks with Guaranteed Truth-Table Generalization

The paper introduces MACCHIATO, a training algorithm that builds a ReLU‑MLP from partial truth‑table data while simultaneously constructing an explicit Boolean circuit over AND, OR, and XOR gates that certifies the network’s computation. The method iteratively projects residuals onto low‑dimensional Boolean classes, compiles the resulting circuit into a ReLU‑MLP, and uses logic minimization and influence‑based variable selection to achieve a six‑layer network with provable truth‑table error bounds. Experiments on synthetic random‑junta tasks show that these certified networks outperform Adam‑trained MLPs in data‑sparse or projection‑aligned regimes and complete faster than flat ESPRESSO in certain settings.

By Hrad Ghoukasian, Anastasis Kratsios
arXiv Machine Learning
Aug 26

Multi-Source Complex Network Reconstruction via Wasserstein Distributionally Robust Optimization and Algorithm Unrolling

The paper introduces MS‑WDRO, a multi‑source Wasserstein distributionally robust optimization framework for reconstructing complex network topologies from scarce target‑domain data and abundant heterogeneous source data. It fuses sources via a weighted Wasserstein barycenter, builds an ambiguity set around it, and solves a regularized Laplacian estimator using a provably convergent ADMM scheme. The authors provide finite‑sample guarantees, demonstrate that naive aggregation is suboptimal, and show through experiments on synthetic data and the ABIDE I neuroimaging dataset that MS‑WDRO outperforms seven baselines in graph recovery, sample efficiency, and diagnostic utility, especially when target samples are limited.

By Chuansen Peng, Yifan Xia, Jinshan Zhong, Xiaojing Shen