arXiv:2606. 11682v1 Announce Type: cross Abstract: Tabular-image multimodal learning aims to improve predictive modeling by jointly using structured tabular attributes and visual data.
By Jiaqi Luo
arXiv:2607. 11007v1 Announce Type: new Abstract: Few-shot multimodal classification commonly attaches a lightweight head, such as $k$-nearest neighbors, logistic regression, or a linear SVM, to a frozen pretrained encoder.
By Jingxiang Zhang, Lujia Zhong, Zijie Zhu, Shuo Huang, Yuang Xu
arXiv:2608. 03557v1 Announce Type: cross Abstract: Tabular-to-image methods that convert tabular data into visual representations have emerged as a novel paradigm for leveraging the high performance of deep learning models.
By Malena Loza, Felipe Grijalva, Eva Milara, Luis Bote-Curiel, Francisco J. Lara-Abelenda, David Chushig-Muzo
The paper introduces TIC‑Bench, a new benchmark for evaluating multimodal large language models on deeply interleaved text‑image contexts. It covers logical, temporal, and spatial association tasks, totaling 2,280 questions across eight specific types. The authors benchmarked ten state‑of‑the‑art MLLMs, finding a significant performance gap versus human experts and highlighting persistent challenges in integrating evidence across interleaved visual and textual inputs.
By Zihao Wang, Xi Xiang, Yuwen Sun, Yingyu Li, Yabo Zhang, Yihan Zeng, Fan Li, Wangmeng Zuo
The paper introduces TEmBed, a unified benchmark for evaluating tabular embeddings across four representation levels—cell, row, column, and table—using a diverse set of models. It demonstrates that the best model depends on the specific task and representation level, providing practical guidance for selecting embeddings in real-world applications. The study aims to facilitate the development of more general-purpose tabular representation models.
By Liane Vogel, Kavitha Srinivas, Niharika D'Souza, Sola Shirai, Oktie Hassanzadeh, Horst Samulowitz
M$^2$PFN is an end‑to‑end multimodal framework that extends the TabPFN in‑context learning engine to Alzheimer’s disease diagnosis by aligning 3D‑MRI and tabular features in a shared subspace. It performs differentiable inference through TabPFN’s transformer, back‑propagates gradients into the encoders, and incorporates a frozen tabular‑only prediction via a gated shortcut. On the ADNI cohort it achieves 65.55 % macro‑F1 and 82.21 % macro‑AUC, surpassing unimodal and multimodal baselines, and it generalizes to external cohorts without retraining.
By Lujia Zhong, Shuo Huang, Jianwei Zhang, Xinyu Nie, Yonggang Shi