arXiv Machine Learning By Kin Ian Lo

The Limits of Binding in Dual Encoders

Read the original on arXiv Machine Learning →

arXiv:2608. 15971v1 Announce Type: new Abstract: Dual-encoder models such as CLIP score an image-caption pair by a single inner product of two independently computed unit vectors, and fail at binding, often scoring near chance when asked to distinguish "a red car and a blue dog" from "a blue car and a red dog".

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Sep 24

Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings

The paper argues that meaning identity—whether two sentences convey the same idea after wording changes—is not encoded in the geometry of independently produced sentence embeddings. Experiments on frozen off‑the‑shelf encoders and language models show that identity can only be reliably computed when both sentences are processed together in a single forward pass, yielding high accuracy (0.90–0.96) on PAWS‑X, whereas independent embeddings or simple fusion methods perform near chance. Even advanced bi‑encoder fine‑tuning improves performance on PAWS but fails to generalize to other similarity tasks, underscoring that identity is a cheap computed operator rather than a property of individual sentence vectors.

By Jiaqi Deng