Most multimodal learning methods improve how heterogeneous representations are aligned and fused, while post-fusion enhancement remains less explored. We propose Parallel Quantum Feature Augmentation (PQFA), a hybrid quantum-classical framework that applies multiple shallow variational quantum circuits to fused multimodal features.
arXiv:2607. 13466v1 Announce Type: new Abstract: Most multimodal learning methods improve how heterogeneous representations are aligned and fused, while post-fusion enhancement remains less explored.
By Mingzhu Wang, Yun Shang
Large-scale Vision-Language Models have demonstrated impressive transfer learning capabilities across a wide range of tasks. For few-shot classification, we observe that VLMs exhibit a notable ability to filter candidate categories and thus achieve high Top-K accuracy.
Mu-DisCoCat is a multimodal variational quantum learning framework that maps Compositional Distributional Semantics (DisCoCat) onto Variational Quantum Circuits (VQCs) to achieve compositional concept generalization (CoCoGen). The pipeline first learns stable object representations from single-object image-text pairs, then fixes these to learn relations in multi-object scenarios. In classical simulations it outperformed a CLIP baseline on out‑of‑distribution relational accuracy, and on noisy quantum emulators and real IBM and IQM hardware it maintained strong fidelity correlations, reliably distinguishing unseen similar and dissimilar pairs.
By Mina Abbaszadeh, Matilda Karabina Moore, Raem Haq, Martha Lewis, Mehrnoosh Sadrzadeh
QiT is a Quantum‑Inspired Transformer designed for visual recognition tasks. It replaces quantum neural network concepts with scalable classical operations: angle‑inspired encoding of image tokens, self‑attention over periodic features approximating quantum fidelity kernels, and gated multiplicative emulation of variational circuit interactions. The model achieves competitive performance on image‑classification benchmarks, matching a classical Transformer while avoiding the high runtime costs of simulated quantum models.
By Badri N. Patro, Vijay Agneeswaran
arXiv:2608. 06846v1 Announce Type: cross Abstract: We test whether a parameterized quantum circuit (PQC) improves a hybrid quantum-classical model's performance on classical datasets, using an interface-matched classical map as the control while holding all other components fixed.
By Hao-Yuan Chen
arXiv:2604. 06135v2 Announce Type: replace-cross Abstract: Efficient data loading remains a bottleneck for near-term quantum machine learning.
By Basil Kyriacou, Viktoria Patapovich, Maniraman Periyasamy, Alexey Melnikov
arXiv:2608. 15601v1 Announce Type: new Abstract: Compositional Concept Generalization (CoCoGen), the ability to systematically recombine learned primitives in novel contexts, is a key challenge for multimodal learning.
By Mina Abbaszadeh, Matilda Karabina Moore, Mehrnoosh Sadrzadeh, Martha Lewis
arXiv:2608. 05595v1 Announce Type: cross Abstract: Circuit cutting lets a large quantum neural network (QNN) run as independent subcircuits on small devices, but rebuilding its outputs by reconstruction carries a classical sampling overhead exponential in the number of cuts - the dominant runtime cost in prior work.
By Prabhjot Singh, Adel N. Toosi, Rajkumar Buyya
arXiv:2608. 04379v1 Announce Type: cross Abstract: We propose a method to optimize the correlation among convolutional neural network (CNN) features that are used as inputs to quantum neural network (QNN) to enhance image classification accuracy.
By Minseo Seong, Youngwook Kim
Circuit cutting lets a large quantum neural network (QNN) run as independent subcircuits on small devices, but rebuilding its outputs by reconstruction carries a classical sampling overhead exponential in the number of cuts - the dominant runtime cost in prior work. We ask whether, for machine-learning tasks, this step is necessary, and replace it with late fusion: each subcircuit is trained and measured independently, and a small classical head combines their outputs - a linear-cost, decision-level combination borrowed from multimodal learning.
arXiv:2603.06755v3 Announce Type: replace
Abstract: We propose a quantum implicit neural representation (QINR)-based autoencoder (AE) and variational autoencoder (VAE) for image reconstruction and ge...
By Saadet M\"uzehher Eren