arXiv Machine Learning By Kerry Luo, Michael Fu, Joshua Peguero, Husnain Malik, Anvay Patil, Joyce Lin, Megan Van Overborg, Ryan Sarmiento, Kevin Zhu

ASCIIBench: Evaluating Language-Model-Based Understanding of Visually-Oriented Text

Read the original on arXiv Machine Learning →

ASCIIBench is a new benchmark that evaluates large language models on generating and classifying ASCII-text images, using a dataset of 5,315 labeled ASCII images. The authors also release a fine‑tuned CLIP model adapted to capture ASCII structure for evaluation. Their analysis shows that cosine similarity on CLIP embeddings fails to separate most categories, indicating a representation bottleneck rather than generational variance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
4d ago

Beyond Bag-of-Words: Diagnosing Compositional Binding Failures in Vision-Language Models

The paper introduces Auto-Comp, a fully automated, concept-driven pipeline that generates photorealistic compositional benchmarks for vision‑language models. Auto‑Comp creates paired Minimal and Contextual samples for each concept, enabling isolation of core binding abilities from visio‑linguistic complexity. Evaluations across 25 models reveal consistent failures in attribute and relational binding, with context helping relational tasks but hindering attribute tasks due to visual clutter.

By Cristian Sbrolli, Toshihiko Yamasaki, Matteo Matteucci
arXiv AI
Aug 26

Giraffe: A Mapping Architecture from Hidden Text Representations to Visual Embeddings for Efficient Graphic Design

The paper introduces Giraffe, a new mapping architecture that converts hidden text token representations into visual embeddings for graphic design tasks. It uses a single [IMG] token per image and two shallow MLP blocks—one for training and one for inference—to compress and expand embeddings, trained with six loss functions. The approach achieves strong performance in both image‑to‑design and text‑to‑design generation while remaining lightweight.

By Nejla Ghaboosi