Neutrality Bites: Gender Representation in AI-Generated Animal Stories
arXiv:2606. 07969v1 Announce Type: cross Abstract: Gender bias in AI-generated stories is a well-documented problem.
arXiv:2606. 07969v1 Announce Type: cross Abstract: Gender bias in AI-generated stories is a well-documented problem.
The paper introduces MObyGaze, a dataset of 20 films annotated by experts for multimodal objectification, covering 6072 segments across 43 hours of video. It defines objectification through a structured thesaurus of 5 sub‑constructs and 11 concepts spanning visual, speech, and audio modalities. The authors formulate learning tasks, explore label diversity strategies, and benchmark vision, text, and audio models to demonstrate the task’s feasibility.
arXiv:2605.31556v2 Announce Type: replace-cross Abstract: Alignment teaches vision-language models (VLMs) to avoid expressing demographic biases, and when gender is clearly visible they largely succe...
arXiv:2606. 09589v1 Announce Type: cross Abstract: AI minidramas (also known as fruit dramas) are short, algorithmically distributed generative AI video series featuring anthropomorphized characters that have recently emerged as a widespread phenomenon on social media platforms.
The paper investigates how in‑context learning (ICL) in large vision‑language models (LVLMs) can amplify gender bias. Using the VL‑BICLE framework, the authors show that gendered ICL demonstrations shift model bias toward the demonstrated gender, especially in tasks involving gendered language such as image captioning and pronoun prediction. They find that similarity‑based retrieval does not mitigate this bias and that replacing real images with synthetic ones from stable diffusion reduces bias without hurting caption quality.
arXiv:2608.21430v1 Announce Type: new Abstract: Multimodal language models increasingly show promise for enabling the large-scale computational analysis of film, opening up new avenues for learning a...
Large language models (LLMs) are increasingly embedded in everyday life and widely used for information seeking, raising concerns about their potential to perpetuate social biases and reinforce stereotypes. In this study, we investigate gender bias in LLMs through the lens of their associations with musical instruments.
Chehre is an emoji‑prompted video dataset designed to study perceptual flexibility in video language models. It contains 2,111 videos of 203 participants expressing 40 facial emojis, with each video annotated by about 30 perceivers, yielding 1,242 annotators in total. The dataset introduces a new task—distributional expression recognition—that evaluates a model’s ability to reproduce the variation seen in human annotations, and shows that persona prompting can shift model perception to better match human variability.
arXiv:2512. 00807v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) inherit significant social biases from their training data, notably in gender representation.
arXiv:2508. 03483v3 Announce Type: replace-cross Abstract: While prior research on text-to-image generation has predominantly focused on biases in human depictions, demographic bias in generated objects remains relatively underexplored.
Narrating the visual component of movies is known as audio description. It is a narrative technique designed to enable blind and visually impaired individuals to follow the story.
The study introduces GAPA, a dataset of 316 physical attributes with 14,706 gender-association ratings from 304 US annotators, showing that such descriptions carry structured gender associations. It evaluates 16 LLMs, finding they partially mirror human ratings but exhibit biases such as compressed distributions, weaker alignment for men, and asymmetric abstention toward non‑binary identities. A proxy model trained on these data is released and applied to analyze character descriptions in LitBank, illustrating the persistence of gendered interpretations in ostensibly neutral language.