Latent Ordinal Evidence, Misaligned Outputs: Inference-Time Ordinal Lens Alignment for Multimodal LLMs
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
Large language models (LLMs) used for ordinal classification exhibit positional bias, where changes in label order, demonstration order, and demonstration placement affect predictions. Systematic experiments across ten frontier LLMs, eight prompt/task/model factors, and five datasets reveal that all models are sensitive to these positional sources, and that accuracy and stability often diverge. Various correction methods, including pointwise, pairwise, and listwise inference, do not reliably mitigate the bias, though a comparison-based listwise approach shows the best overall balance yet varies across models and bias types.
ReWEIGH the Evidence is a training‑free decoding technique that calibrates token‑level ordinal visual evidence to reduce hallucinations in large vision‑language models. It aggregates vocabulary ranks across visual positions, compares candidates to a token‑specific reference derived from unlabeled images, and applies a bounded penalty only when evidence falls below this reference. Experiments on four 7B backbones show up to a 21.3% reduction in hallucinated object mentions while largely preserving or improving descriptive and general performance, with minimal added latency.
arXiv:2607. 23575v1 Announce Type: cross Abstract: Ordinal regression is widely used in scenarios where labels are discrete yet inherently ordered.
arXiv:2607. 08109v1 Announce Type: new Abstract: We propose contrastive order learning (ConOrd), a contrastive learning framework for ordinal regression that integrates the strengths of contrastive learning and order learning.
Lumen is a pathology vision‑language model that aligns frozen unimodal foundation models (Virchow2 and BioMedBERT) using rank‑4 adapters and projection heads, training only 0.40% of the total parameters on the QUILT‑1M corpus. It achieves the highest mean chance‑corrected balanced accuracy (0.546) across nine zero‑shot patch benchmarks and demonstrates strong performance on lymph‑node metastasis detection, with AUROC scores of 0.964 internally and 0.955 externally. While it ranks third in cross‑modal retrieval, Lumen’s low‑parameter training yields competitive results at both patch and slide levels.
arXiv:2608. 09227v1 Announce Type: new Abstract: Omnimodal language models (OLMs) enable unified audio-visual understanding, but processing long joint token sequences makes inference computationally prohibitive.