arXiv Machine Learning

MSNN-LINet: Cross-Modal Learning via Continuous Linear Integration

arXiv:2606. 31135v1 Announce Type: cross Abstract: We present LINet (Linear Integration Network), a Multi-Stream Neural Network (MSNN) for RGB-D scene classification.

arXiv Machine Learning
Jul 14

Vertical Fusion: Condensing Internal Representations for Robust ViT Classification

arXiv:2607. 10391v1 Announce Type: cross Abstract: Despite exposing rich intermediate representations, Vision Transformers (ViTs) are almost exclusively utilized as black-box feature extractors, where only the last layer is considered for downstream tasks.

By Francesco Di Salvo, Shyam Nandan Rai, Hamed Damirchi, Ignacio Meza De la Jara, Sebastian Doerrich, Marco Lents, Christian Ledig
Hugging Face Trending Papers
Jul 22

Current Injection Spiking Neural Network for Infrared and Visible Image Fusion

Infrared and visible image fusion (IVIF) integrates the complementary information of two modalities into a single image with richer scene content. While existing methods are largely built on artificial neural networks (ANNs), which densely compute over all activations, spiking neural networks (SNNs) communicate through sparse binary spikes and compute only where and when a spike occurs, offering a route to more energy-efficient fusion.

arXiv Machine Learning
Aug 27

CAT-GS: Balanced Multimodal Learning via Calibrated Gating and Fusion Surgery

The paper introduces CAT‑GS, a training controller that stabilizes multimodal neural networks by addressing three failure modes: modality imbalance, unstable gating, and fusion interference. CAT‑GS calibrates teacher-derived reliability, applies a margin‑thresholded gating policy, caps gradient budgets, and uses fusion‑only PCGrad, all without altering model architectures or losses. Experiments on audio‑visual, tri‑modal, synthetic, and cross‑domain benchmarks show that CAT‑GS matches or surpasses strong imbalance‑aware baselines while producing smoother gating and fewer conflicting fusion gradients.

By Mahir Shahriar Tamim, Sharjil Khan, Md. Samiul Alim, Tanvir Ahmed Khan, Shafin Rahman, Nabeel Mohammed
Hugging Face Trending Papers
Jul 2

DRDN: Decoupled Representation Dynamic Network for From-Scratch ViT Class-Incremental Learning

Dynamic expansion methods for class-incremental learning (CIL) protect task-specific knowledge by growing dedicated tokens or subnetworks, yet our analyses suggest that classification supervision alone does not sufficiently preserve task-agnostic shared backbone representations over long incremental sequences. We identify two intertwined challenges: cross-task confusion from sequential training on predominantly current-task data, which biases decision boundaries toward recent tasks; and under-optimized shared representations in the backbone that cap long-term discriminability as tasks accumulate.

arXiv AI
Jul 21

Emergent Hierarchical Monosemantic Neurons from the Group-Contrastive Forward-Forward Algorithm

arXiv:2607. 16295v1 Announce Type: cross Abstract: Mechanistic interpretability has made significant strides in understanding neural network representations, with sparse dictionary learning (SDL) methods, most prominently sparse autoencoders, as a central paradigm.

By Yiming Tang, Qinglin Qi, Zhaoqian Yao, Harshvardhan Saini, Dianbo Liu