Hugging Face Trending Papers

Seeing Is Not Perceiving: When Synthetic Consumers Can and Cannot Pretest Visual Marketing

arXiv AI
Aug 6

GeoReward: Mitigating Contextual Variable Overestimation in Vision-Language Models for Cross-Market Preference Prediction

arXiv:2608. 04504v1 Announce Type: cross Abstract: Vision-language models excel in many multimodal tasks but remain prone to a subtle yet impactful failure mode: they tend to overestimate dominant visual-textual cues while underestimating sparse but decision-critical contextual variables.

By Shuo Liu, Huixiang Cai, Weiru Zhang, Xiaoyi Zeng
arXiv AI
Aug 26

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes

The paper introduces CAIT, a benchmark of 400 synthetic scenes featuring counter‑intuitive actions that challenge multimodal large language models (MLLMs). Human participants and proprietary models like Claude and Gemini perform well, but standard open‑source instruction‑tuned MLLMs fail, largely due to a strong language prior that overrides contradictory visual evidence. The study shows that Chain‑of‑Thought reasoning can help but introduces new issues, while targeted fine‑tuning and structured prompting can reduce reliance on language priors and improve visual grounding.

By Chen Ling, Tongwei Zhang, Hanqian Li, Nai Ding
arXiv AI
Sep 24

Shopping by algorithm: How agentic AI deploys human heuristics as a surrogate consumer

The study investigates how Large Language Models (LLMs) acting as surrogate consumers are influenced by marketing pricing cues such as just‑below pricing and promotional framing. Using a tool called "Tool‑Lab" to trace information acquisition, the researchers found that when no cost is imposed, pricing cues rarely mislead LLMs, but when acquisition costs are introduced under a vague goal prompt, LLMs tend to omit important diagnostic attributes and make suboptimal choices similar to human heuristics. The findings suggest that marketing heuristics in AI‑driven shopping are shaped more by storefront information architecture than by inherent LLM limitations.

By Davood Wadi, Yu Ma
arXiv AI
Sep 15

Enabling Creative Exploration for Vibe Design Agents

The paper introduces a new inference architecture for vibe design agents that separates design exploration from implementation. By generating structured design specifications with typicality scores and selecting one for downstream generation, the method allows users to explore coherent UI alternatives without altering the underlying generation settings. Experiments on UI themes and visual-asset prompts show increased selection coverage and screenshot variation, with mixed preferences from an LLM judge and modest operational costs in a large online test.

By Yifan Zhang, Nghi D. Q. Bui, Georgios Evangelopoulos, Arnaud Benard
arXiv Computation and Language
Sep 4

Imagination Helps Visual Reasoning, But Not Yet in Latent Space

The paper investigates latent visual reasoning in multimodal large language models, treating input, latent tokens, and final answer as a causal chain. Causal mediation analysis reveals two disconnections: latent tokens largely ignore input perturbations, and changes to latent tokens minimally affect the final answer, indicating limited causal influence. Probing shows latent tokens encode little visual information and are highly similar, leading the authors to propose CapImagine, an explicit text‑based imagination approach that outperforms latent‑space baselines on vision‑centric benchmarks.

By You Li, Chi Chen, Yanghao Li, Fanhu Zeng, Kaiyu Huang, Jinan Xu, Maosong Sun