Most multimodal learning methods improve how heterogeneous representations are aligned and fused, while post-fusion enhancement remains less explored. We propose Parallel Quantum Feature Augmentation (PQFA), a hybrid quantum-classical framework that applies multiple shallow variational quantum circuits to fused multimodal features.
arXiv:2607. 13466v1 Announce Type: new Abstract: Most multimodal learning methods improve how heterogeneous representations are aligned and fused, while post-fusion enhancement remains less explored.
By Mingzhu Wang, Yun Shang
arXiv:2608. 15601v1 Announce Type: new Abstract: Compositional Concept Generalization (CoCoGen), the ability to systematically recombine learned primitives in novel contexts, is a key challenge for multimodal learning.
By Mina Abbaszadeh, Matilda Karabina Moore, Mehrnoosh Sadrzadeh, Martha Lewis
The paper introduces Quantum-Inspired Nonlinear Adapters (QINA), compact modules that apply learnable trigonometric feature lifting followed by bounded nonlinear aggregation to pretrained vision models. QINA enables structured oscillatory basis functions with a norm-dependent Lipschitz bound, allowing spectral reshaping of representations without expanding the receptive field or significantly increasing parameters. Experiments on natural and medical imaging tasks show that QINA consistently outperforms identity baselines, fixed Fourier mappings, and parameter-matched generic adapters, demonstrating that geometry- and spectrum-aware adaptation is crucial for effective frozen-backbone transfer learning.
By Mostafa Mehdipour Ghazi
QiT is a Quantum‑Inspired Transformer designed for visual recognition tasks. It replaces quantum neural network concepts with scalable classical operations: angle‑inspired encoding of image tokens, self‑attention over periodic features approximating quantum fidelity kernels, and gated multiplicative emulation of variational circuit interactions. The model achieves competitive performance on image‑classification benchmarks, matching a classical Transformer while avoiding the high runtime costs of simulated quantum models.
By Badri N. Patro, Vijay Agneeswaran
arXiv:2606. 02785v1 Announce Type: new Abstract: Large machine learning models benefit substantially from multimodal inputs that provide a complementary view of the same example.
By Aritra Bal, Michael Binder, Markus Klute, Benedikt Maier, Michael Spannowsky