arXiv:2608. 11661v1 Announce Type: cross Abstract: A multiplicative dual-encoder network computes a real-valued output for a pair of inputs as the inner product of their separate encodings.
By Zijian Zhao, Sen Li
arXiv:2606. 03879v1 Announce Type: cross Abstract: As foundation models scale toward fusing more heterogeneous visual streams, understanding how diverse encoders interact under joint training becomes a prerequisite for principled design.
By Wei Ding, Yudong Zhang, Ruobing Xie, Xingwu Sun, Jiansheng Chen, Yu Wang
arXiv:2603. 01568v2 Announce Type: replace Abstract: Efficient coding theory predicts that biological perceptual systems compress sensory input optimally under resource constraints, with the systematic structure of errors reflecting the geometry of that compression.
By Leyla Roksan Caglar, Pedro A. M. Mediano, Baihan Lin
The paper introduces Action-Contrastive Masked Transition Modeling (AC‑MTM), a method that replaces the Gaussian regularizer used in Joint‑Embedding Predictive Architectures (JEPAs) with a contrastive inverse‑dynamics head. AC‑MTM trains a forward latent‑prediction model while an auxiliary inverse‑dynamics task forces the encoder to distinguish actions from latent transitions, preventing collapse without requiring a target network or reconstruction loss. Experiments on pixel‑control and multi‑object visual tasks show that AC‑MTM matches or surpasses the performance of the Gaussian‑based SIGReg regularizer, achieving up to 20–24 point improvements on the OGBench Visual Scene benchmark.
By Jack Boylan, Chris Hokamp
arXiv:2609.08561v1 Announce Type: new
Abstract: Class disentanglement (the separation of a representation's class-conditional point clouds along depth and over training) is usually read off descripti...
By Sushovan Majhi
The paper introduces the Intersection Euler Characteristic Profile, a topological metric for measuring class overlap in neural representations, and provides exact permutation and sign‑flip tests to assess disentanglement across layers. Using this statistic, the authors analyze 111 networks and 52,650 measurements, finding that disentanglement is depth‑graded, occurs early, and is influenced by training choices such as augmentation and weight decay. The study also demonstrates that the unnormalized mass of the profile predicts test accuracy, while the dimensionless quotient does not outperform simple linear probes.
By Sushovan Majhi
The study evaluates a frozen Geneformer representation for predicting CRISPRi perturbation effects under a tightly controlled, pre‑registered protocol. While the representation shows significant predictive power within the Virtual Cell Challenge dataset, it fails to transfer to external screens, with negative zero‑shot Spearman correlations. The analysis also reveals that the VCC endpoint is heavily influenced by sampling depth, as cell count alone explains most of the variance, indicating a sampling‑depth entanglement that could mask transfer failures in less controlled settings.
By Mehrdad Shoeibi, Niloofar Yousefi
The paper investigates how knowledge distillation from Vision Transformers to smaller CNNs can cause dimensional collapse in the student’s representation space. Using SVD and Shannon entropy, the authors show that cosine‑based distillation leads to a drastic reduction in effective rank, while adding an InfoNCE objective can double the rank but harms downstream accuracy due to signal dilution. They further demonstrate that a label‑aware contrastive objective (Supervised Contrastive distillation) can maintain or improve accuracy without unnecessary rank expansion, indicating that effective rank alone is not a reliable indicator of representation quality.
By Kabir Thayani
arXiv:2608. 20054v1 Announce Type: new Abstract: Multi-module systems often expose every module to the full input.
By Narcis Marincat
EXPL-FR is a lightweight adapter that aligns a vision‑language model’s image encoder with a frozen face‑recognition (FR) embedding space, enabling the FR model to be explained using semantic attribute prompts without any text training. By mapping 978 attribute prompts across 22 categories into the FR space, the method identifies the most detectable concepts—forming a readable semantic signature that better separates identities than the full vocabulary. The approach is evaluated on four FR backbones and two VLM encoders, providing identity‑level, per‑image, and differential explanations, and demonstrates that prompt‑driven audits can rank FR models by per‑ethnicity error and attribute‑change verification cost without requiring labeled data.
By Guray Ozgur, Mustafa Efe Tamyapar, Naser Damer, Fadi Boutros
The paper introduces Action-Contrastive Masked Transition Modeling (AC‑MTM), a method that stabilizes Joint‑Embedding Predictive Architectures (JEPAs) without relying on Gaussian regularization. AC‑MTM adds a training‑only inverse‑dynamics head that uses Action‑NCE to force each latent transition to identify its generating action, thereby preventing encoder collapse. Experiments on pixel‑control and multi‑object visual tasks show that AC‑MTM trains stably from scratch and matches or surpasses the performance of SIGReg, achieving up to a 24‑point improvement on the OGBench Visual Scene benchmark.
arXiv:2607. 16292v4 Announce Type: replace-cross Abstract: Brain-encoding foundation models predict fMRI responses to video, audio and text well enough to win the Algonauts 2025 challenge.
By Carson Rodrigues