arXiv AI By Kordel Kade France

COLIP-2: Olfaction-Vision-Language Embeddings

Read the original on arXiv AI →

arXiv:2607. 17559v1 Announce Type: cross Abstract: The Contrastive Olfaction-Language-Image Pre-training 2 (COLIP-2) model is a multimodal embeddings space that places olfaction as a first-class citizen among vision and language.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Jul 29

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation

arXiv:2607. 25527v1 Announce Type: cross Abstract: Unifying visual understanding and generation in one model holds immense promise, but remains challenging and expensive due to heavy compute and data demands and conflicts between the visual features needed for these two capabilities.

By Weiming Zhuang, Jiabo Huang, Jingtao Li, Zhizhong Li, Chen Chen, Sina Sajadmanesh, Lingjuan Lyu
arXiv Machine Learning
Aug 7

Dynamic Object Masks as Goal Representations for Visual Goal-Conditioned Reinforcement Learning

arXiv:2510. 06277v2 Announce Type: replace-cross Abstract: Goal-conditioned reinforcement learning (GCRL) offers a unified way to pursue diverse tasks, yet most existing methods rely on state- or position-based goal representations that are unavailable in real-world robotics.

By Fahim Shahriar, Cheryl Wang, Alireza Azimi, Gautham Vasan, Hany Hamed, Abhishek Naik, A. Rupam Mahmood, Colin Bellinger
Hugging Face Trending Papers
Jun 22

Humanoid-OmniOcc: Stereo-Based Full-View Occupancy Dataset for Embodied AI

Occupancy prediction at voxel-level granularity is essential for safe robotic navigation and interaction in complex environments. Existing occupancy datasets, however, are predominantly designed for autonomous driving with vehicle-centric biases -- forward-facing cameras, far-field geometry, and static road priors -- limiting their applicability to embodied humanoid perception.

arXiv AI
Jul 9

GemNav: Discrete-Token Visual Robot Navigation using a Multimodal Large Language Model

arXiv:2607. 06882v1 Announce Type: cross Abstract: Visual navigation policies built on large pretrained models have so far followed a common recipe: a dedicated visual encoder, a bespoke action head, and training on thousands of hours of cross-embodiment datasets.

By Peter Bohm, Saimunur Rahman, Abdelwahed Khamis, Sagun Man Singh Shrestha, Chris McCool, Peyman Moghadam