Sparse Autoencoders for Interpretable Out-of-Distribution Detection
arXiv:2607. 12094v1 Announce Type: cross Abstract: Reliable detection of out-of-distribution (OOD) samples is crucial for the safe deployment of machine learning models.
The paper demonstrates that incorporating code automorphisms into syndrome-based neural decoding (SBND) improves the models’ learning and generalization through data augmentation during training and inference. By applying this technique to short, high-rate codes, the authors achieve performance close to maximum likelihood decoding (MLD) using small datasets and appropriate training. The study also indicates that previous SBND results may have underestimated their true error‑correction capability due to insufficient training.
arXiv:2607. 12094v1 Announce Type: cross Abstract: Reliable detection of out-of-distribution (OOD) samples is crucial for the safe deployment of machine learning models.
arXiv:2608. 15412v1 Announce Type: cross Abstract: Encoder-based code representation models remain widely deployed for discriminative tasks such as clone detection and code classification, where their small size and low inference cost are decisive.
arXiv:2509.10452v3 Announce Type: replace-cross Abstract: Pretrained automatic speech recognition (ASR) models such as Whisper perform well but still need domain adaptation to handle unseen parlance....
The paper investigates using synthetic natural-language descriptions to contrastively pretrain small transformer encoders for code representation. By pairing generated descriptions with code in a dual-encoder setup during training and discarding them at inference, the authors achieve significant improvements over traditional pretraining baselines on most evaluated tasks. When fine‑tuned, these models match or surpass much larger zero‑shot models and remain competitive with execution‑aware supervision, indicating a scalable alternative for code embeddings.
arXiv:2603. 04198v2 Announce Type: replace-cross Abstract: Sparse autoencoders (SAEs) are widely used to extract human-interpretable features from neural network activations, but their learned features can vary substantially across random seeds and training choices.
arXiv:2608.22956v1 Announce Type: new Abstract: Concept-bottleneck controllable generation routes multi-attribute control through a low-dimensional concept code that, at deployment, must be synthesis...
arXiv:2606. 16246v1 Announce Type: cross Abstract: As AI labs approach a data ceiling where compute capacity outpaces the rate of new high-quality text generation, language model pretraining is shifting toward a data-constrained, compute-abundant regime that demands productive multi-epoch training on fixed corpora.
arXiv:2607. 04011v1 Announce Type: cross Abstract: While decoders have rapidly scaled, encoders have remained largely unchanged since BERT.
arXiv:2510. 07884v2 Announce Type: replace-cross Abstract: Weak-to-strong generalization provides a promising paradigm for scaling large language models (LLMs) by training stronger models on samples from aligned weaker ones, without requiring human feedback or explicit reward modeling.
The paper introduces Dynamic DAE Guardrails (DSG), a method that uses Dynamic Sparse Autoencoders to perform precision unlearning in large language models. DSG leverages principled feature selection and a dynamic classifier to target activation-based unlearning, outperforming existing gradient‑based methods in terms of computational efficiency, stability, sequential unlearning, resistance to relearning attacks, data efficiency, and interpretability.
arXiv:2508. 17320v3 Announce Type: replace Abstract: Understanding the internal representations of large language models (LLMs) remains a central challenge for interpretability research.
The paper introduces the Superposed Latent Autoencoder (SLAE), a method that stores multiple wide latent representations together by superposing them into a single memory tensor using learned codes and randomized keys. SLAE eliminates the need for tight dimensional bottlenecks, achieving up to 56% lower reconstruction error on datasets such as CIFAR-10/100 and SVHN while maintaining the same storage budget. The approach also boosts downstream classification performance by up to 16.79 percentage points, demonstrating that wide representations can be effectively compressed through structured interference rather than dimensional reduction.