Representable but Unlearned: Encoding Rank and the Interaction-Prediction Floor
Read the original on arXiv AI →The paper investigates how input encodings constrain the set of contrasts a predictor can reproduce, even when no individual contrast is forced to zero. By computing the attainable contrast space from an encoder’s equivalence classes and a fixed contrast design—without using labels, loss, or a fitted model—the authors derive an empirical error floor for any unrestricted decoder on those classes. Experiments on a 140‑rectangle siRNA interaction panel show that a graph neural network’s training‑only feature mask reduces the rank of interaction contrasts from 140 to 72, creating a floor of 0.009980 (14.6% of the fitted model’s interaction squared error). Removing the mask eliminates the floor but only marginally improves MSE, while restoring chemistry columns recovers full rank. A separate RNA‑splicing predictor with an injective encoding achieves full rank and a zero floor, illustrating that the encoding itself, not the model, limits recoverable contrast space.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.