arXiv Machine Learning

Text Has Curvature

The paper investigates whether natural language text possesses an intrinsic curvature, proposing a new metric called Texture that captures word-level discrete curvature. Texture is defined as a signed two-axis curvature of the word-in-context belief field, measuring how context from one side contracts or expands the semantic effect of context from the other side. The authors provide empirical and theoretical evidence of non-flat semantic inference, define Texture formally, and demonstrate its practical utility in improving long-context inference and retrieval-augmented generation.

arXiv Computer Vision
Aug 31

GraSP-VL: Length as a Semantic Granularity Interface for Vision-Language Representations

GraSP‑VL demonstrates that the length of frozen vision‑language embeddings can serve as a controllable semantic interface. By learning a shared near‑orthogonal prefix transform, the method creates a Semantic Matryoshka where short prefixes encode coarse semantics and longer prefixes reveal finer language‑grounded distinctions, all while preserving the original embedding geometry. Experiments on COCO/Flickr30K and SugarCrepe‑clean show strong performance with negligible drift in the full embedding space.

By Zesheng Li, Chengchang Pan, Honggang Qi
arXiv AI
Jun 24

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models

arXiv:2605. 08245v4 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) increasingly power high-stakes applications, from medical imaging to autonomous systems, yet they routinely hallucinate, confidently describing content not present in the input.

By Harshvardhan Saini, Samyak Jha, Yiming Tang, Dianbo Liu
arXiv Computer Vision
Sep 3

Deeply Interleaved Text-Image Contexts for Multimodal LLMs Assessment

The paper introduces TIC‑Bench, a new benchmark for evaluating multimodal large language models on deeply interleaved text‑image contexts. It covers logical, temporal, and spatial association tasks, totaling 2,280 questions across eight specific types. The authors benchmarked ten state‑of‑the‑art MLLMs, finding a significant performance gap versus human experts and highlighting persistent challenges in integrating evidence across interleaved visual and textual inputs.

By Zihao Wang, Xi Xiang, Yuwen Sun, Yingyu Li, Yabo Zhang, Yihan Zeng, Fan Li, Wangmeng Zuo
arXiv AI
Sep 7

Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models

The article presents a technical manual for an open toolkit designed to measure how transformer language models individuate word meanings across different contexts. It introduces the concept of a "bridge form"—a single word that appears unchanged in multiple domains but with distinct senses—and outlines a full pipeline from specifying these forms to extracting layer-wise representations, computing silhouette-based separation metrics, and visualizing results. The manual details each design choice and its intended methodological safeguards, emphasizing that it serves as a methodological reference rather than reporting empirical findings.

By Jos\'e Luciano Ver\c{c}osa Marques, Frederico Jorge Heitmann, Daniel Omar Perez, Marcelo Vinicius de Paula, T\'arcio Andr\'e dos Santos Barros
arXiv Machine Learning
Sep 24

ASCIIBench: Evaluating Language-Model-Based Understanding of Visually-Oriented Text

ASCIIBench is a new benchmark that evaluates large language models on generating and classifying ASCII-text images, using a dataset of 5,315 labeled ASCII images. The authors also release a fine‑tuned CLIP model adapted to capture ASCII structure for evaluation. Their analysis shows that cosine similarity on CLIP embeddings fails to separate most categories, indicating a representation bottleneck rather than generational variance.

By Kerry Luo, Michael Fu, Joshua Peguero, Husnain Malik, Anvay Patil, Joyce Lin, Megan Van Overborg, Ryan Sarmiento, Kevin Zhu