Most multimodal learning methods improve how heterogeneous representations are aligned and fused, while post-fusion enhancement remains less explored. We propose Parallel Quantum Feature Augmentation (PQFA), a hybrid quantum-classical framework that applies multiple shallow variational quantum circuits to fused multimodal features.
arXiv:2607. 13466v1 Announce Type: new Abstract: Most multimodal learning methods improve how heterogeneous representations are aligned and fused, while post-fusion enhancement remains less explored.
By Mingzhu Wang, Yun Shang
arXiv:2608. 15601v1 Announce Type: new Abstract: Compositional Concept Generalization (CoCoGen), the ability to systematically recombine learned primitives in novel contexts, is a key challenge for multimodal learning.
By Mina Abbaszadeh, Matilda Karabina Moore, Mehrnoosh Sadrzadeh, Martha Lewis
The paper introduces Quantum-Inspired Nonlinear Adapters (QINA), compact modules that apply learnable trigonometric feature lifting followed by bounded nonlinear aggregation to pretrained vision models. QINA enables structured oscillatory basis functions with a norm-dependent Lipschitz bound, allowing spectral reshaping of representations without expanding the receptive field or significantly increasing parameters. Experiments on natural and medical imaging tasks show that QINA consistently outperforms identity baselines, fixed Fourier mappings, and parameter-matched generic adapters, demonstrating that geometry- and spectrum-aware adaptation is crucial for effective frozen-backbone transfer learning.
By Mostafa Mehdipour Ghazi
QiT is a Quantum‑Inspired Transformer designed for visual recognition tasks. It replaces quantum neural network concepts with scalable classical operations: angle‑inspired encoding of image tokens, self‑attention over periodic features approximating quantum fidelity kernels, and gated multiplicative emulation of variational circuit interactions. The model achieves competitive performance on image‑classification benchmarks, matching a classical Transformer while avoiding the high runtime costs of simulated quantum models.
By Badri N. Patro, Vijay Agneeswaran
arXiv:2606. 02785v1 Announce Type: new Abstract: Large machine learning models benefit substantially from multimodal inputs that provide a complementary view of the same example.
By Aritra Bal, Michael Binder, Markus Klute, Benedikt Maier, Michael Spannowsky
arXiv:2608. 06846v1 Announce Type: cross Abstract: We test whether a parameterized quantum circuit (PQC) improves a hybrid quantum-classical model's performance on classical datasets, using an interface-matched classical map as the control while holding all other components fixed.
By Hao-Yuan Chen
arXiv:2608. 04379v1 Announce Type: cross Abstract: We propose a method to optimize the correlation among convolutional neural network (CNN) features that are used as inputs to quantum neural network (QNN) to enhance image classification accuracy.
By Minseo Seong, Youngwook Kim
arXiv:2604. 06135v2 Announce Type: replace-cross Abstract: Efficient data loading remains a bottleneck for near-term quantum machine learning.
By Basil Kyriacou, Viktoria Patapovich, Maniraman Periyasamy, Alexey Melnikov
arXiv:2605. 27923v2 Announce Type: replace-cross Abstract: The rapid growth of computer vision and increasingly complex image recognition tasks has exposed fundamental computational limitations of classical machine learning models, motivating the exploration of quantum computing as an emerging new paradigm.
By Sudip Vhaduri, Ryan Gammon, Sayanton Dibbo
arXiv:2506. 01850v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable success in instruction-following tasks by integrating pretrained visual encoders with large language models (LLMs).
By Wayner Barrios, Andr\'es Villa, Juan Le\'on Alc\'azar, SouYoung Jin, Bernard Ghanem
We propose a method to optimize the correlation among convolutional neural network (CNN) features that are used as inputs to quantum neural network (QNN) to enhance image classification accuracy. Unlike prior approaches that employ orthogonal decomposition as preprocessing, we intentionally introduce correlated features that are more physically compatible with QNN.