SemiScope: Disentangling Classifier Tuning and Joint Optimization in Semi-Supervised Security Classification
arXiv:2607. 00113v1 Announce Type: new Abstract: Background.
arXiv:2607. 00113v1 Announce Type: new Abstract: Background.
The paper introduces C-Score, a diagnostic framework for evaluating pseudo‑label‑based semi‑supervised learning (SSL) when unlabeled data may contain out‑of‑distribution (OOD) samples. C-Score assesses training behavior across prediction, feature representation, and optimization, using metrics such as PLE, CCI, Sem‑Drift, and Grad‑Align. Experiments on CIFAR‑10 and CIFAR‑100 with various OOD sources show that C‑Score detects hidden degradation that clean accuracy alone fails to reveal, highlighting the need for internal diagnostic signals in SSL robustness assessment.
arXiv:2605. 28418v3 Announce Type: replace Abstract: With the rise of tabular foundation models alongside traditional models still performing well on many tasks, choosing the right model for a tabular dataset remains difficult.
The paper introduces a pre‑training pipeline that creates transformer‑based imputation specialists for tabular data with specific missingness patterns. By featurizing entries, generating synthetic data with configurable missingness modules, and fitting on millions of synthetic tables, the pipeline produces pattern‑specific models that outperform dedicated methods for each missingness pattern. A default model trained only on MCAR data, TabImpute, remains robust across all tested patterns, and the authors release the pipeline, models, and a new benchmark of 42 datasets and 11 missingness patterns.
arXiv:2609.26839v1 Announce Type: cross Abstract: Post-hoc probability calibration is usually evaluated under an optimistic assumption: the held-out calibration labels are clean. In many AI deploymen...
arXiv:2607. 00958v1 Announce Type: new Abstract: Time series are central to modern data mining applications, from industrial telemetry and server metrics to finance and physiology, yet time-series self-supervised learning often depends on view and augmentation choices that encode domain-specific invariances.
GEAR is a two‑stage framework that distills tabular foundation models into lightweight MLP or tree‑based predictors for efficient CPU deployment. In the first stage, synthetic covariates are used as teacher‑query locations to train the student on soft TFM targets, expanding coverage beyond observed rows. The second stage re‑anchors the student to the target distribution using real labels and out‑of‑fold teacher predictions, preventing self‑labeling leakage and improving performance. Experiments on TALENT and TabArena show that GEAR‑distilled MLPs outperform supervised MLPs by up to 2.00 AUC points on binary tasks and 1.35 on multiclass tasks, and also outperform CatBoost, while dramatically reducing inference time and memory usage.
arXiv:2407. 05370v3 Announce Type: replace Abstract: Semi-supervised learning (SSL) algorithms often struggle to perform well when trained on imbalanced data.
arXiv:2607. 19847v1 Announce Type: cross Abstract: Predicting missing cell values in tabular data is a fundamental problem in data cleaning.
arXiv:2607. 16681v1 Announce Type: new Abstract: Early sepsis prediction from electronic health records is challenged by irregular sampling, high missingness, and class imbalance.
arXiv:2606. 30410v1 Announce Type: cross Abstract: Foundation models for predictive machine learning on tabular data have recently gained significant traction in academia and industry.
arXiv:2605. 12705v2 Announce Type: replace Abstract: How can we train models whose post-trained capabilities survive subsequent fine-tuning?