arXiv AI

The Label Imitation Game: Turing Test Network for Zero-Shot Pseudo-Label Pruning

arXiv:2606. 30875v1 Announce Type: cross Abstract: Foundation model pseudo-labeling - labeling data strictly via zero-shot inference - enables massive scale, but performance is undermined by hallucinations that evade standard thresholds.

arXiv AI
Sep 25

Self-Play Pretraining with Zero Data

The paper introduces Self‑Play Pretraining with Zero Data, a proof‑of‑concept method that lets a model generate its own training data by searching over all computable processes using a universal Turing machine. Two models— a generator that proposes byte‑sequence programs and a learner that predicts those sequences—train together, with the generator rewarded for producing data at the learner’s frontier, creating an adaptive curriculum. Experiments show that zero‑shot performance on natural datasets scales predictably with compute, and the models exhibit in‑context learning and discover mathematical sequences during training.

By Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine
arXiv AI
Aug 6

Adversarially Robust Abductive Fusion of Pre-trained Transformer-based Perception Models

arXiv:2608. 04190v1 Announce Type: new Abstract: Deploying pre-trained perception models in novel environments degrades their accuracy under distributional shift, and assembling them alone does not recover it: combiners such as majority voting trade recall for precision and are brittle to coordinated failures.

By Mario Leiva, Yue Ma, Qinru Qiu, Gerardo Simari, Paulo Shakarian
arXiv AI
Jun 15

Learning What to Predict: Downstream-Guided Task Design for Continued Pretraining

arXiv:2601. 22108v2 Announce Type: replace-cross Abstract: Continued pretraining is optimized with fixed self-supervised tasks but selected by downstream performance, creating a coarse feedback loop in which practitioners evaluate checkpoints, change data mixtures or objectives, and restart runs, while individual updates remain blind to target capabilities.

By Shuqi Ke, Giulia Fanti