Persona Prompting in Multimodal Urban Perception: Descriptive Convergence and Interpretive Variation
Read the original on arXiv Computation and Language →The paper investigates how persona prompting influences the language produced by two multimodal large language models—Qwen3‑VL and Gemma4—when describing urban scenes. Outputs are categorized into descriptive grounding (captions), perception tags, and interpretive framing (justifications). Results show captions largely converge across persona profiles, while justifications differ markedly, especially along economic status, political orientation, and personality dimensions, with economic status producing the greatest variation. Perception tags also reflect attribute similarities, and exploratory topic analysis indicates persona‑specific evaluative emphasis. Overall, persona prompting has a stronger effect on interpretive framing than on descriptive grounding.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.