Cheap and Powerful Tests for Supervised Subspaces: Per-Component Inference for PLS
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2608. 02412v1 Announce Type: new Abstract: Large language models (LLMs) have become the default tool for a remarkable range of tasks, yet they have had conspicuously little success at one of the most common machine learning workloads: predictive analytics over tabular data.
The article introduces PTED, a Python implementation of a permutation test based on the Energy Distance for two-sample testing in multiple dimensions. PTED uses pairwise distances to compute a test statistic that works in high dimensions, on learned feature representations, and for any data type where a distance can be defined. The authors demonstrate that PTED scales linearly with dimensions and sample size while retaining strong discriminative power, and show it outperforms other multi‑dimensional tests in sensitivity.
arXiv:2609.39263v1 Announce Type: new Abstract: A concept subspace's effect on model behavior does not establish how it relates to the output readout. We introduce a two-sided geometric diagnostic th...
arXiv:2511. 11041v2 Announce Type: replace-cross Abstract: We find that current sentence-embedding models produce outputs with a consistent bias: every embedding $e$ decomposes as $\tilde e + \mu$, where the mean $\mu$ is near-identical across all sentences.
arXiv:2608. 05238v1 Announce Type: new Abstract: Training multimodal models to align time series with language runs into a self-supervision trap.
arXiv:2606. 18011v1 Announce Type: cross Abstract: Constraint-based causal discovery relies on repeated conditional independence tests, but fast nonparametric tests often sacrifice calibration, especially when variables depend on the conditioning set through nonlinear relationships.