Quantum Entangled Multimodal Fusion Networks (QEMFN): Resource-Aware Hybrid Vision-Language Fusion via Trainable Entanglement
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
Most multimodal learning methods improve how heterogeneous representations are aligned and fused, while post-fusion enhancement remains less explored. We propose Parallel Quantum Feature Augmentation (PQFA), a hybrid quantum-classical framework that applies multiple shallow variational quantum circuits to fused multimodal features.
arXiv:2607. 13466v1 Announce Type: new Abstract: Most multimodal learning methods improve how heterogeneous representations are aligned and fused, while post-fusion enhancement remains less explored.
Large-scale Vision-Language Models have demonstrated impressive transfer learning capabilities across a wide range of tasks. For few-shot classification, we observe that VLMs exhibit a notable ability to filter candidate categories and thus achieve high Top-K accuracy.
Mu-DisCoCat is a multimodal variational quantum learning framework that maps Compositional Distributional Semantics (DisCoCat) onto Variational Quantum Circuits (VQCs) to achieve compositional concept generalization (CoCoGen). The pipeline first learns stable object representations from single-object image-text pairs, then fixes these to learn relations in multi-object scenarios. In classical simulations it outperformed a CLIP baseline on out‑of‑distribution relational accuracy, and on noisy quantum emulators and real IBM and IQM hardware it maintained strong fidelity correlations, reliably distinguishing unseen similar and dissimilar pairs.
QiT is a Quantum‑Inspired Transformer designed for visual recognition tasks. It replaces quantum neural network concepts with scalable classical operations: angle‑inspired encoding of image tokens, self‑attention over periodic features approximating quantum fidelity kernels, and gated multiplicative emulation of variational circuit interactions. The model achieves competitive performance on image‑classification benchmarks, matching a classical Transformer while avoiding the high runtime costs of simulated quantum models.
arXiv:2608. 06846v1 Announce Type: cross Abstract: We test whether a parameterized quantum circuit (PQC) improves a hybrid quantum-classical model's performance on classical datasets, using an interface-matched classical map as the control while holding all other components fixed.