Frozen encoders are chosen by how well a lightweight head reads a finding from their features, not whether the geometry separates it. Nearest-neighbor discordance does, but with unequal banks the opposite-label neighbor wins on density, not geometry, so prevalence alone makes an uninformed encoder look blind.
arXiv:2607.18451v3 Announce Type: replace
Abstract: A foundation encoder is pretrained once on a large image corpus and then reused with its weights frozen. Each new task is solved by training a smal...
By Soroosh Tayebi Arasteh, Sven Nebelung, Daniel Truhn
arXiv:2608. 16709v1 Announce Type: cross Abstract: A radiologist reading a model's output faces two problems.
By Vignesh Nagarajan, Sriram Venkatapathy
The paper critiques current memorization audits for generative models, arguing that lacking a proper null distribution leads to misleading conclusions. It introduces two exact null tests—one permutation test for whole models and a calibrated test for single images—showing that many previously flagged memorizations disappear under these stricter controls. The authors also propose a scale‑restricted statistic based on the Intersection Euler Characteristic Profile to better detect distinct copied images.
By Sushovan Majhi, Pramita Bagchi
arXiv:2604.11508v3 Announce Type: replace
Abstract: Fine-tuning a pretrained classifier leaves some samples reliably learned and others cycling between correct and incorrect. Curriculum learning, dat...
By Miit Daga, Swarna Priya Ramu
The paper demonstrates that open‑ended Theory‑of‑Mind trackers can produce valid beliefs that are absent from finite reference sets, and that treating unmatched outputs as false can reverse model‑selection rankings. By recoding references for 259 beliefs, the authors show a dramatic drop in weighted prevalence and a reversal of strictly proper Brier risk, with similar distortions observed in a 301‑question NQ‑open DPR‑BERT pipeline. The study further reveals that 90‑96% of audited unmatched beliefs are literally true, and introduces a TriSource‑Restore method that anchors reference labels to a probability‑sampled human pilot to restore calibration and ranking integrity.
By Zhexi Feng, Wuxi Chen, Bingrui Zhang