arXiv Machine Learning

The Importance of Encoder Choice:A Tabular-Image Study

arXiv:2607. 07756v1 Announce Type: new Abstract: Multimodal learning usually requires a dedicated encoder per modality.

arXiv AI
Jul 29

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation

arXiv:2607. 25527v1 Announce Type: cross Abstract: Unifying visual understanding and generation in one model holds immense promise, but remains challenging and expensive due to heavy compute and data demands and conflicts between the visual features needed for these two capabilities.

By Weiming Zhuang, Jiabo Huang, Jingtao Li, Zhizhong Li, Chen Chen, Sina Sajadmanesh, Lingjuan Lyu
arXiv Machine Learning
Jun 4

Towards Pretraining Text Encoders for TabPFN

arXiv:2606. 04876v1 Announce Type: new Abstract: Tabular foundation models, such as TabPFN, achieve strong performance on tabular datasets with numerical and categorical data, but do not natively handle high-cardinality text features.

By Mustafa Tajjar, Alexander Pfefferle, Lennart Purucker, Frank Hutter
arXiv AI
Jun 2

LLMs Need Encoders for Semantic IDs Too

arXiv:2606. 00324v1 Announce Type: cross Abstract: Multimodal LLMs use dedicated encoders to bridge non-language modalities (vision encoders for images, depth models for audio codec tokens) because raw token embeddings alone cannot capture modality-specific structure.

By Xiangyi Chen, Zelun Wang, Xinyi Li, Yi-Ping Hsu, Jaewon Yang, Jiajing Xu