arXiv Machine Learning

A machine-learning-assisted progressive digit-randomness screening framework for detecting non-random patterns in raw numerical research data

arXiv:2606. 07128v1 Announce Type: new Abstract: Raw numerical datasets remain less systematically examined in integrity screening than images, plagiarism, or summary-statistic inconsistencies.

arXiv AI
Jul 24

Synthetic minority data is redundant or invalid: a data-dependent validity theory and a de-biased test

arXiv:2607. 20787v1 Announce Type: cross Abstract: For two decades, the standard remedy for class-imbalanced learning has been to fabricate synthetic minority examples, and the standard evidence of their validity has been a check that cannot fail: synthetic points are scored against the very data that generated them.

By Ahmad B. Hassanat, Ahmad S. Tarawneh, Ghada A. Altarawneh