arXiv Machine Learning

Parameter-Efficient Adapter Tuning for Tabular-Image Multimodal Learning

arXiv:2606. 11682v1 Announce Type: cross Abstract: Tabular-image multimodal learning aims to improve predictive modeling by jointly using structured tabular attributes and visual data.

arXiv AI
4d ago

AdaKerNet: Neural Kernel Decoding for Task-Adaptive Prediction with Multimodal Large Models

AdaKerNet is a task‑adaptive neural kernel decoder that operates on frozen multimodal representations from large foundation models, without requiring access to the models’ parameters. It learns Lipschitz‑controlled multimodal features, a reference kernel providing a soft structural prior, and a lightweight nonlinear predictor that deforms this structure. Experiments on four multimodal large language models and diverse input modalities show consistent improvements over baseline decoders, achieving up to 41% error reduction in scarce‑label settings.

By Konstantinos D. Polyzos, Eleni Oikonomou, Tara Javidi
arXiv Machine Learning
Aug 21

Table2Image: Lightweight Tabular Learning with Generated Proxy Representations and Reliability Diagnostics

arXiv:2412. 06265v3 Announce Type: replace Abstract: Deep tabular models should ideally balance predictive performance, parameter efficiency, and robustness to imperfect learning signals---properties that are rarely considered jointly.

By Seungeun Lee, Kihwan Lee, Subin Bae, Sangjun Lee, Seulbin Lee, Julia Stoyanovich, Il-Youp Kwak, Seungsang Oh
arXiv Machine Learning
Sep 23

MMAP: Multimodal Missing-Aware Pretraining for Longitudinal Alzheimer's Prediction

MMAP is a Multimodal Missing‑Aware Alignment Pretraining method designed to learn image‑tabular representations from incomplete data. It uses a sigmoid contrastive learning image encoder with generative reconstruction, a tabular encoder based on a foundation model, and a missing token generator to handle missing modalities. The approach is evaluated on longitudinal Alzheimer’s tasks—predicting disease stage conversion and amyloid status—and outperforms both multimodal and unimodal baselines.

By Fiona Kekwick, Matthew Baugh, Bernhard Kainz, Paul M. Matthews, Wenjia Bai
arXiv Computer Vision
Sep 22

Enabling Vision and Cross-Modal Learning for Multimodal Stroke Recurrence Prediction: An Interpretable Two-Step Framework

arXiv:2609.22271v1 Announce Type: new Abstract: Multimodal stroke recurrence prediction requires effective integration of heterogeneous clinical and imaging data, yet modality imbalance often causes...

By Christian Gapp, Elias Tappeiner, Martin Welk, Karl Fritscher, Stephanie Mangesius, Constantin Eisenschink, Philipp Deisl, Michael Knoflach, Astrid E. Grams, Elke R. Gizewski, Rainer Schubert
arXiv Computer Vision
Sep 16

FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation

FLAT (Flexible‑Length Aligned Transmodal representations) is a joint multimodal pre‑training framework that learns a shared encoder for images and text, producing 1‑D continuous embeddings that can be directly used by downstream generative decoders. By combining contrastive alignment with bidirectional cross‑modal generative objectives, FLAT yields representations that are both discriminative and generative, enabling cross‑modal retrieval and generation with a single pre‑training stage. The model achieves strong performance on T2I generation (GenEval 71.1), image captioning (BLEU‑4 40.5, CIDEr 138.6), and retrieval tasks (Recall@5 86.8/75.8 on MS‑COCO, 98.3/93.6 on Flickr30K), and supports linear interpolation, latent space arithmetic, and zero‑shot composed retrieval.

By Guangyu Sun, Shlok Kumar Mishra, Wentao Bao, Robert Zhenheng Yang, Xiao Wang, Xiyuan Wang, Yujunrong Ma, Chen Yuan, Max Xiangjun Fan, Jun Xiao, Jianpeng Cheng