Towards High-Level Semantic Intelligence
arXiv:2607. 24082v1 Announce Type: new Abstract: Recent advances in AI have substantially expanded its cognitive and reasoning capabilities.
Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.
arXiv:2607. 24082v1 Announce Type: new Abstract: Recent advances in AI have substantially expanded its cognitive and reasoning capabilities.
arXiv:2607. 22794v1 Announce Type: cross Abstract: Automatic depression detection with deep learning has shown promise but often suffers from limited generalization due to domain shift arising from inter-speaker variability.
arXiv:2607. 22864v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) excel at visual interpretation but fail on spatial reasoning tasks that humans solve reliably.
arXiv:2607. 22600v1 Announce Type: new Abstract: Information visualizations are widely used to communicate patterns, trends, and outliers, yet deceptive design choices-such as truncated or inverted axes, distorted aspect ratios, inappropriate encodings, and misleading color mappings-can systematically alter interpretation while preserving the underlying data.
arXiv:2607. 23514v1 Announce Type: cross Abstract: Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence.
arXiv:2607. 22999v1 Announce Type: cross Abstract: Language agents can now interact fluently with users in software, but robots still struggle to bring comparable interaction to physical tasks.
arXiv:2607. 23909v1 Announce Type: new Abstract: Many recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as the action backbone.
arXiv:2607. 22721v1 Announce Type: cross Abstract: Cognitive remediation tasks often require patients to perform structured actions involving object manipulation and sequential reasoning.
arXiv:2607. 23258v1 Announce Type: new Abstract: Large-scale calcium imaging has created an opportunity to build foundation-style models for neural population dynamics, but a central question remains unresolved: \textbf{whether a model pretrained on one collection of recordings can generalize to new datasets, experimental paradigms, and even species.
arXiv:2607. 22632v1 Announce Type: new Abstract: The rapid rise of vlogs as a personalized storytelling medium has created a demand for automated systems to evaluate and refine vlog editing plans.
arXiv:2607. 23913v1 Announce Type: new Abstract: Modern vision-language models (VLMs) increasingly rely on dynamic or high-resolution visual encoding, producing thousands of visual tokens that substantially increase downstream language-model inference cost.
arXiv:2607. 23944v1 Announce Type: new Abstract: Human visual reasoning typically follows a coarse-to-fine attention process, starting from global scene understanding and gradually focusing on question-relevant regions.
arXiv:2607. 22586v1 Announce Type: new Abstract: Key-Value (KV) caching is essential for efficient inference in multimodal large language models (MLLMs), yet its memory footprint grows linearly with context length and becomes a major bottleneck due to the large number of visual tokens.
arXiv:2607. 24683v1 Announce Type: cross Abstract: Multi-modal classification leverages complementary information across diverse data sources to enhance predictive performance.
arXiv:2607. 22689v1 Announce Type: new Abstract: Graphical user interface (GUI) agents are systems powered by large multimodal models (LMMs).
arXiv:2607. 22643v1 Announce Type: new Abstract: Multimodal retrieval-augmented generation (mRAG) aims to answer image-text queries with external knowledge, but most existing systems still retrieve directly from raw multimodal input over a flat evidence space.
arXiv:2607. 22712v1 Announce Type: cross Abstract: Single-cell light microscopy images have become an important data source for characterizing cell phenotypes, but their complexity and heterogeneity pose challenges to high-throughput automated analysis.
arXiv:2607. 23245v1 Announce Type: cross Abstract: Multimodal Federated Learning is often challenged by arbitrary modality missingness and Non-IID data distributions, which lead to severe representation drift and hinder effective collaboration across clients.
arXiv:2607. 24259v1 Announce Type: new Abstract: Recent work has explored the use of Large Language Models (LLMs) to automate simulation model building, typically by generating executable code directly from natural language descriptions.
arXiv:2607. 23910v1 Announce Type: cross Abstract: Cooperative perception through vehicle-to-everything (V2X) communication can overcome the inherent physical limitations of individual autonomous vehicles, such as occlusions and limited sensor range.