Large-scale Vision-Language Models have demonstrated impressive transfer learning capabilities across a wide range of tasks. For few-shot classification, we observe that VLMs exhibit a notable ability to filter candidate categories and thus achieve high Top-K accuracy.
Most multimodal learning methods improve how heterogeneous representations are aligned and fused, while post-fusion enhancement remains less explored. We propose Parallel Quantum Feature Augmentation (PQFA), a hybrid quantum-classical framework that applies multiple shallow variational quantum circuits to fused multimodal features.
arXiv:2607. 13466v1 Announce Type: new Abstract: Most multimodal learning methods improve how heterogeneous representations are aligned and fused, while post-fusion enhancement remains less explored.
By Mingzhu Wang, Yun Shang
arXiv:2603.06755v3 Announce Type: replace
Abstract: We propose a quantum implicit neural representation (QINR)-based autoencoder (AE) and variational autoencoder (VAE) for image reconstruction and ge...
By Saadet M\"uzehher Eren
arXiv:2606. 02785v1 Announce Type: new Abstract: Large machine learning models benefit substantially from multimodal inputs that provide a complementary view of the same example.
By Aritra Bal, Michael Binder, Markus Klute, Benedikt Maier, Michael Spannowsky
arXiv:2608. 11884v1 Announce Type: cross Abstract: Quantum generative adversarial networks (QGANs) have attracted increasing attention for image generation using parameterized quantum circuits.
By Xue Yang, Rigui Zhou, ShiZheng Jia, Dax Enshan Koh, Siong Thye Goh, Young-Wook Cho, YaoChong Li, Xuezhi Ma, Hongyu Chen, Xin Wang
Quantum generative adversarial networks (QGANs) have attracted increasing attention for image generation using parameterized quantum circuits. Existing amplitude-based approaches face two key limitations: pixel locations are typically encoded by computational-basis indices or address qubits, causing quantum resources to grow with image resolution; meanwhile, jointly decoding many pixels from normalized quantum states introduces probability competition among pixels and limits precise pixel-wise control.
The paper introduces Quantum-Inspired Nonlinear Adapters (QINA), compact modules that apply learnable trigonometric feature lifting followed by bounded nonlinear aggregation to pretrained vision models. QINA enables structured oscillatory basis functions with a norm-dependent Lipschitz bound, allowing spectral reshaping of representations without expanding the receptive field or significantly increasing parameters. Experiments on natural and medical imaging tasks show that QINA consistently outperforms identity baselines, fixed Fourier mappings, and parameter-matched generic adapters, demonstrating that geometry- and spectrum-aware adaptation is crucial for effective frozen-backbone transfer learning.
By Mostafa Mehdipour Ghazi
arXiv:2503.24111v4 Announce Type: replace-cross
Abstract: Graph Neural Networks (QGNNs) offer a promising approach to combining quantum computing with graph-structured data processing. While classica...
By Arthur M. Faria, Ignacio F. Gra\~na, Savvas Varsamopoulos
arXiv:2510. 02528v2 Announce Type: replace Abstract: Large Multimodal Models (LMMs) demonstrate impressive in-context learning abilities from few multimodal demonstrations, yet the internal mechanisms supporting such task learning remain opaque.
By Shuhao Fu, Esther Goldberg, Ying Nian Wu, Hongjing Lu
arXiv:2604. 06135v2 Announce Type: replace-cross Abstract: Efficient data loading remains a bottleneck for near-term quantum machine learning.
By Basil Kyriacou, Viktoria Patapovich, Maniraman Periyasamy, Alexey Melnikov
UVU is a vision-language unified autoregressive framework that integrates visual supervision directly into the pre-training stage of multimodal large language models. By using continuous visual encoding and a large-scale iterative hierarchical clustering algorithm to build a pixel-level visual codebook, UVU enables lossless representation of visual inputs and autoregressive generation of pixel-level image tokens alongside textual tokens. This approach synergizes pixel-level visual perception with semantic-level visual understanding, allowing models to internalize visual reconstruction capabilities and improve multimodal understanding performance.
By Zhehan Kan, Xinghua Jiang, Yubo Zhu, Yanlin Liu, Xiaochen Yang, Zhixiang Wei, Shifeng Liu, Qingmin Liao, Wenming Yang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun