arXiv AI

Deep Learning Models Also Recall Features

The paper discusses how large language models retrieve facts from their weights, proposing that this phenomenon reflects a broader operation termed feature recall. It argues that a linear projection can be interpreted as retrieving stored information scaled by input activations, and demonstrates that feature recall applies across various architectures, contrasting it with the traditional feature combination paradigm. The authors also explore potential mechanistic identification of feature recall cases and suggest new empirical directions for mechanistic interpretability research.

Hugging Face Trending Papers
Jul 24

Same Predictions, Different Reasons: The Effect of Quantization on Model Explanations

Post-training quantization (PTQ) has become a practical solution for deploying deep learning models on resource-constrained edge devices by compressing high-precision floating-point weights into low-precision representations without requiring retraining. Past research has demonstrated that quantization largely preserves classification accuracy; however, whether it also preserves the model's internal reasoning remains an open question.

arXiv Machine Learning
Jul 28

Same Predictions, Different Reasons: The Effect of Quantization on Model Explanations

arXiv:2607. 22872v1 Announce Type: new Abstract: Post-training quantization (PTQ) has become a practical solution for deploying deep learning models on resource-constrained edge devices by compressing high-precision floating-point weights into low-precision representations without requiring retraining.

By Kazi Kamruzzaman Rabbi, Md. Zami Al Zunaed Farabe, M. Sohel Rahman