Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

4,626 stories · RSS feed

arXiv AI
Jun 10

MetaPlate: Counterfactual-Guided RAG-LLM Tool for Personalized Food Recommendation and Hyperglycemia Prevention

arXiv:2606. 10120v1 Announce Type: cross Abstract: Postprandial hyperglycemia is a key risk factor for metabolic disorders; however, existing dietary guidance is often static, impractical, and insufficiently personalized, providing recommendations that are difficult to follow or not impactful.

By Asiful Arefeen, Carol Johnston, Hassan Ghasemzadeh
arXiv AI
Jun 10

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering

arXiv:2510. 04514v3 Announce Type: replace Abstract: Recent multimodal LLMs have shown promise in chart-based visual question answering, but their performance declines sharply on unannotated charts-those requiring precise visual interpretation rather than relying on textual shortcuts.

By Rachneet Kaur, Nishan Srishankar, Zhen Zeng, Sumitra Ganesh, Manuela Veloso
arXiv Machine Learning
Jun 10

Spatiotemporal Graph Transformer for 3D Neighborhood Interaction and Quality Prediction in Metal Additive Manufacturing

arXiv:2606. 10227v1 Announce Type: new Abstract: Metal additive manufacturing enables the fabrication of complex parts, but achieving consistent build quality remains challenging due to interactions induced by repeated layer-wise melting, solidification, and reheating across the 3D build.

By Joyce Karen Pelaez, Siqi Zhang, Hoo Sang Ko
arXiv AI
Jun 10

Soul Computing: A Theoretical Framework and Technical Architecture for Intelligent Agents with Independent Consciousness

arXiv:2606. 10413v1 Announce Type: new Abstract: Breakthroughs in large language models and multimodal generation technologies have propelled the digital reconstruction of human mental traits, emotional patterns, and long-term memory from science fiction toward engineering practice.

By Jinshan Zhang, Xishi Zhou, Qiu Peng, Jianwei Yin
arXiv AI
Jun 10

A History-Aware Visually Grounded Critic for Computer Use Agents

arXiv:2606. 11078v1 Announce Type: new Abstract: Various test-time interventions for Computer Use Agents (CUAs), including critic models, have been developed to improve performance through pre-execution action evaluation in complex Graphical User Interface (GUI) environments.

By Jaewoo Lee, Zaid Khan, Archiki Prasad, Justin Chih-Yao Chen, Supriyo Chakraborty, Kartik Balasubramaniam, Sambit Sahu, Elias Stengel-Eskin, Hyunji Lee, Mohit Bansal