arXiv AI

Cheap Probes Predict Expensive Training in 3D-CT Vision--Language Models

arXiv:2607. 22771v1 Announce Type: cross Abstract: Picking the frozen image encoder for a 3D~CT vision--language model (VLM), together with the token-compression scheme on top of it, is a search over many candidates.

arXiv AI
Jul 24

Sparse Concept Channels in Frozen 3D CT Vision Encoders

arXiv:2607. 20993v1 Announce Type: cross Abstract: Large vision-language models are becoming increasingly dominant in 3D medical image interpretation, but we rarely know which internal units encode clinical findings or where that information lives in the representation.

By Farhad Nooralahzadeh, Lea Bogensperger, Christian Bluethgen, Michael Krauthammer
arXiv Computation and Language
Sep 4

Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs

The study investigates how long‑video language models decide which frames to keep, compress, and reuse, testing each decision in isolation across six selection rules, three benchmarks, and two answering models. It finds that selecting frames based on queries yields the biggest performance boost, that halving spatial resolution costs little, and that reallocating saved tokens to more compressed frames can further improve accuracy. The work also highlights the importance of a unified evaluation harness to avoid misleading comparisons.

By Prakhar Khatri
arXiv AI
Aug 11

Resolution Meets Reduction: Efficient Visual Context for 3D Radiology Report Generation

arXiv:2608. 08713v1 Announce Type: cross Abstract: Vision-language models offer a promising path toward automating radiology report generation, but applying them to full 3D CT volumes poses substantial computational challenges.

By Jonathan Suprijadi, Raphael Stock, Moritz Langenberg, David Zimmerer, Kim-Celine Kahl, Stefan Denner, Yannick Kirchhoff, Karol Gotkowski, Maximilian Rokuss, Jeremias Traub, Tassilo Wald, Constantin Ulrich, Klaus Maier-Hein
arXiv Computer Vision
Sep 18

SCOUT: Sim-to-Real Text-Based Person Retrieval by Embedding-Space Prediction over Frozen Video Features

SCOUT is a frozen‑encoder approach for sim‑to‑real text‑based person retrieval that predicts cross‑modal embeddings instead of fine‑tuning cross‑encoders. It uses a trainable predictor to map patch tokens from a frozen video encoder (V‑JEPA) into the embedding space of a frozen text encoder (EmbeddingGemma), guided by a bidirectional InfoNCE objective. The method achieves state‑of‑the‑art results on the AI City Challenge 2026 Track 4, with a full retrieve‑fuse‑rerank pipeline reaching 84.25 mAP@10 and a single frozen model alone scoring 60.63, while training costs are modest (≈95 GPU‑hours).

By Abdarahmane Traor\'e, Andy Couturier, \'Eric Hervet
arXiv AI
Aug 20

ReWEIGH the Evidence: Calibrating Token-Level Ordinal Visual Evidence to Mitigate Hallucinations in Large Vision-Language Models

ReWEIGH the Evidence is a training‑free decoding technique that calibrates token‑level ordinal visual evidence to reduce hallucinations in large vision‑language models. It aggregates vocabulary ranks across visual positions, compares candidates to a token‑specific reference derived from unlabeled images, and applies a bounded penalty only when evidence falls below this reference. Experiments on four 7B backbones show up to a 21.3% reduction in hallucinated object mentions while largely preserving or improving descriptive and general performance, with minimal added latency.

By Jihae Jeong, Junha Choi, Hwanjo Yu