Cross-Modal Emotion Understanding: A Transformer-GAT Approach for Dialogue Emotion Recognition proposes a hybrid framework that combines a Transformer and a Graph Attention Network to capture both global semantic information and fine-grained relationships between modalities. The model is evaluated on the IEMOCAP and MELD datasets, achieving weighted F1 scores of 72.45% and 77.37%, respectively, and surpasses state‑of‑the‑art methods. These results suggest that integrating multimodal features with balanced global and local context modeling can provide deeper emotional insights for dialogue emotion recognition.
By Jiaqi Qiao, Yifan Lyu, Xiujuan Xu
TraceGuard is an adaptive, rank‑based filtering method designed to detect and remove poisoned image‑text pairs from multimodal training corpora. It evaluates six corpus‑level features—such as cross‑modal neighborhoods, recurring text, and changes after text‑span erasure—to identify suspicious examples without training the victim model. In experiments across 19 attack scenarios, TraceGuard successfully removes an average of 98.4% of poisoned data while discarding only 5.4% of clean data, reducing residual attack impact to at most 1% in most configurations.
By Haoyang Li, Yaxin Xiao, Linyan Dai, Jiawen Fu, Zi Liang, Jason Xue, Qingqing Ye, Haibo Hu
OncoVision is a privileged‑information training framework that learns from mammography images and clinical data during training but performs inference using only mammographic images. It employs an attention‑based encoder‑decoder to jointly segment masses, calcifications, axillary findings, and breast tissue, and predicts ten structured clinical features such as BI‑RADS. Two late‑fusion strategies (Independent and Dependent) integrate imaging, radiomic, and clinical information to improve diagnostic precision, and a retrospective multi‑reader study showed higher diagnostic confidence, reduced reading time, and segmentation accuracy comparable to or better than radiologists.
By Istiak Ahmed, Galib Ahmed, K. Shahriar Sanjid, Md. Tanzim Hossain, Md. Nishan Khan, Md. Misbah Khan, Md. Arifur Rahman, Sheikh Anisul Haque, Sharmin Akhtar Rupa, Mohammed Mejbahuddin Mia, Mahmud Hasan Mostofa Kamal, Md. Mostafa Kamal Sarker, M. Monir Uddin
Distributed learning in embodied reinforcement-learning agents offers a degree of privacy by retaining raw sensor data on-device and transmitting only policy gradients to the server. Yet temporal stru...
OmniFabric is a new method for creating high‑quality, globally coherent texture maps for 3D garment reconstruction from a single image. It first generates a coarse texture initialization on the garment’s sewing pattern using a 3D mesh and Vision‑Language Model priors, then refines this in the UV domain with a diffusion transformer conditioned on 3D positional features. The approach removes distortion and baked‑in artifacts, producing photorealistic 3D garments that outperform existing baselines.
The paper investigates whether layer-wise visual‑text similarity in multimodal large language models (MLLMs) truly reflects cross‑modal content integration. By injecting Gaussian noise into the visual stream of 13 MLLMs, the authors show that common scalar alignment metrics (CKA, SVCCA, MIR, principal‑angle cosine) fail to distinguish corrupted from intact visual tokens, a phenomenon they term the "alignment illusion." They propose the principal‑angle gap (PA gap) as a more reliable geometric diagnostic that correlates better with task accuracy and reveals when internal geometry diverges from performance.
M3GD is a novel approach for robotic novel view synthesis that fuses camera images and LiDAR point clouds without requiring a separate cross‑modal translator. By projecting LiDAR data onto the image latent grid and injecting it via a lightweight residual adapter, M3GD enhances both RGB and depth generation on the GrandTour dataset compared to image‑only baselines. Experiments on a ground robot confirm that the method can be deployed in real‑world scenarios with a tunable quality‑cost trade‑off.
Agentic-GER is an LLM-based agent designed to correct terminology in long‑form speech transcripts. It leverages global context from the full transcript to flag suspicious terms, selectively re‑transcribes the source audio to verify candidate corrections, and uses accepted edits to inform future decisions. Experiments on GigaSpeechBench with four LLMs and two ASR systems show consistent terminology improvements in both Chinese and English, achieving up to a 36.8% relative reduction in biased character error rate over the Whisper baseline.
The paper introduces Domain Recentering with Confidence Calibration (DRC), a training‑free technique that adapts CLIP to unlabeled target images by fitting a Gaussian mixture and applying posterior‑weighted mean subtraction, followed by a log‑prior correction based on confidence‑weighted predictions. DRC improves cross‑domain accuracy, surpassing zero‑shot CLIP by 4.13 and 5.07 points on ViT‑B/16 and ResNet‑50, and maintains gains under ImageNet distribution shifts.
The paper introduces PixelJev, a native-image decision interface that combines an image, a task instruction, and a runtime candidate set to produce a structured choice and candidate-conditioned probabilities using small open multimodal models. It unifies recognition and multiple-choice visual question answering via a language-model readout, offering options for frozen inference, language-side adaptation, and held-out calibration. Across seven benchmarks, 64-shot source adaptation significantly boosts Pets accuracy from 60.13% to 92.40%, and the model supports VQA tasks without target fitting, though calibration and cross-family transfer remain challenges.
The paper shows that chain‑of‑thought (CoT) instructions can distort multiple‑choice vision‑language model evaluation when a scorer appends a reasoning cue but reads answer‑label logits before the model generates any rationale. This CoT‑prefix scoring causes significant drops in accuracy (e.g., Qwen2.5‑VL‑7B falls from 80.76% to 45.48% on ScienceQA) and leads most predictions to choose the first option. Analysis reveals that while answer information remains linearly accessible in late layers, the immediate readout is misled by probability mass shifting toward continuation tokens, and the issue varies across datasets and models.
Spot, Separate, and Enhance (SSE) is the first multimodal, user‑guided generative model for audio remixing and enhancement. It rebalances audio, removes unwanted sources, and reduces reverberation in video content, guided by both video and textual descriptions. The authors introduce the DegradedMix dataset, built on MuddyMix, and use generative‑model evaluation metrics to demonstrate SSE’s superior controllability and remixing quality compared to existing baselines.
The paper introduces EASE, an encoder‑only model for Audio‑Visual Semantic Segmentation that eliminates redundant components found in prior Transformer‑based approaches. EASE achieves state‑of‑the‑art accuracy while running at up to 365 FPS—three times faster than previous models—and can be trained in under 11 GPU‑hours. The authors provide code, weights, and samples, positioning EASE as a scalable foundation for future research and real‑time applications.
RGBD20K is a new large-scale RGB‑D dataset designed to advance semantic segmentation research. It contains 20,000 image pairs annotated with 160 fine‑grained categories, far exceeding the diversity of existing benchmarks such as NYUv2 and SUN RGB‑D. The authors also provide high‑fidelity annotations and introduce a score‑purified fusion (SPF) method that achieves state‑of‑the‑art results on multiple benchmarks.
The paper introduces Latent Action Driving Annotations (LADA), a three‑stage pipeline that converts large amounts of unlabelled observation‑trajectory data into a language‑conditioned driving model. First, a latent action model with a vector‑quantised bottleneck learns a compact codebook of vehicle intents. Then, a small set of language‑annotated examples trains a vision‑language translator to map observations and instructions into this codebook, and finally a VLA is trained on observation‑latent‑action pairs across the full corpus. Using less than 5% of language annotations, LADA attains a Driving Score of 87.98 and a Success Rate of 70.46% on Bench2Drive, matching or surpassing fully supervised baselines.
By Alexey Zakharov, Kemal Oksuz, Puneet K. Dokania
The paper presents a dual‑encoder Transformer model for estimating Planetary Boundary Layer Height (PBLH) from satellite radiances, addressing challenges of multimodal, spatially incomplete data. It benchmarks eight different approaches, analyzes model reliance via grouped Shapley decomposition, and demonstrates that the proposed architecture achieves a mean absolute error of 155.8 m on a global test set, outperforming all baselines. On out‑of‑distribution data from the TEAMx campaign, the model attains 165.3 m MAE, better than a pixel‑wise baseline trained on the same data.
By Lorenzo Innocenti, Luca Catalano, Edoardo Arnaudo, Claudio Rossi, Salvatore Larosa, Domenico Cimini, Paolo Garza
ChronoSteer is a decoupled agentic framework that bridges large language models and time series foundation models by learning cross‑modal alignment from synthetic paired supervision. It converts textual events into revision instructions that steer a frozen time‑series model, discretizes these instructions into a compact codebook to reduce semantic divergence, and then refines the predictions with a two‑stage training strategy. The authors also release a leakage‑controlled multimodal benchmark and report a 25.8% improvement in zero‑shot prediction accuracy over the unimodal backbone.
By Chengsen Wang, Qi Qi, Zhongwen Rao, Lujia Pan, Jingyu Wang
ASCIIBench is a new benchmark that evaluates large language models on generating and classifying ASCII-text images, using a dataset of 5,315 labeled ASCII images. The authors also release a fine‑tuned CLIP model adapted to capture ASCII structure for evaluation. Their analysis shows that cosine similarity on CLIP embeddings fails to separate most categories, indicating a representation bottleneck rather than generational variance.
By Kerry Luo, Michael Fu, Joshua Peguero, Husnain Malik, Anvay Patil, Joyce Lin, Megan Van Overborg, Ryan Sarmiento, Kevin Zhu
PRISM‑VLM is a new benchmark for compact vision‑language models that evaluates each item across seven axes—task quality, behavioral robustness, and capability bottlenecks—rather than collapsing performance into a single accuracy score. It aggregates these axes into a single PScore while also providing per‑axis profiles, revealing differences that single‑axis benchmarks miss, such as sycophancy. The benchmark draws on items from fifteen public datasets and will be released with its full pipeline, prompts, and annotations.
By Sanghee Park, Kee-Eung Kim