arXiv AI
Sep 24

Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms

The paper introduces a generation benchmark for culturally specific kinship terms, evaluating five open‑weight large language models (LLMs) on Hindi, Tamil, and Korean. Unlike prior multiple‑choice tests that treat kinship understanding as a recognition task, the study prompts LLMs to generate terms across two communicative tasks and compares results to a matched option‑supported baseline. Findings show that while models like GPT‑OSS120B and Llama‑3.370B can select correct terms in over 90% and 78% of cases respectively, they produce the correct term only 36% and 24% of the time, indicating a significant evaluation‑format gap and highlighting the difficulty of culturally specific kinship generation even when relationships are explicitly stated.

By Sahil Pardasani, Madhusudan Singh