← Back to all news
arXiv Computer Vision September 22, 2026 By Kiran Naseer, Samreen Azhar, Dwarikanath Mahapatra

Reassessing Global Gradient-Norm Imbalance in BLIP Fine-Tuning Across Physical Domains

Read the original on arXiv Computer Vision →

The Flow has not summarised this story yet — read it at arXiv Computer Vision.

  • llms
  • fine-tuning
  • multimodal

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Computer Vision
Sep 10

UOT-Gap: A Variational Principle for the Modality Gap in Vision-Language Models via Unbalanced Optimal Transport

arXiv:2609.10224v1 Announce Type: new Abstract: Vision-language models such as CLIP embed images and text in a shared space, where modality-specific distributions often remain separated. Existing acc...

By Zonglin Yang, Huilan Ma, Xudan Zheng, Yuejun Xie
llmsragmultimodalsafety
More like this →
arXiv AI
Jul 13

GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs

arXiv:2507. 18043v2 Announce Type: replace-cross Abstract: Inference-time steering methods offer a lightweight alternative to fine-tuning large language models (LLMs) and vision-language models (VLMs) by modifying internal activations at test time without updating model weights.

By Duy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal
llmsfine-tuningmultimodalsafety
More like this →
arXiv Machine Learning
6d ago

Not All Layers Need Tuning: Diagnosing and Directing Adaptation in Vision-Language-Action Models

arXiv:2609.18084v1 Announce Type: cross Abstract: Fine-tuning a Vision-Language-Action (VLA) model for a new deployment environment is expensive, yet most methods apply uniform-capacity adapters to e...

By Shahram Najam Syed, Arthur Jakobsson, Prayuj Sachdev, Jeffrey Ichnowski
fine-tuningmultimodalsafety
More like this →
arXiv Machine Learning
Sep 2

The Visual Insensitivity Gap: Diagnosing When Vision-Language Models Fail to Use Visual Evidence

arXiv:2609.00868v1 Announce Type: cross Abstract: Vision-language models are evaluated by aggregate accuracy on multimodal benchmarks, a practice that implicitly assumes the model uses its visual inp...

By Genpei Zhang
llmsmultimodalbenchmarks
More like this →
arXiv Machine Learning
Jun 9

Steer Where It Matters: Token-Level Visual-Sensitivity Steering for LVLMs Hallucination Mitigation

arXiv:2606. 07647v1 Announce Type: cross Abstract: Large vision language models (LVLMs) have made rapid advancements and are deployed across various applications, yet hallucinations remain a major challenge.

By Ruipeng Zhang, Zhihao Li, C. L. Philip Chen, Tong Zhang
llmsmultimodalbenchmarkssafety
More like this →
arXiv Computer Vision
Sep 16

HyCal: A Training-Free Prototype Calibration Method for Cross-Discipline Few-Shot Class-Incremental Learning

arXiv:2604.15678v2 Announce Type: replace Abstract: Pretrained Vision-Language Models (VLMs) like CLIP show promise in continual learning, but existing Few-Shot Class-Incremental Learning (FSCIL) met...

By Eunju Lee, MiHyeon Kim, JuneHyoung Kwon, Yoonji Lee, JiHyun Kim, Soojin Jang, YoungBin Kim
llmsragmultimodalbenchmarkssafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea