The Concrete-Arbitrary Gap: Kinship Reasoning in LLMs Is Not Indifferent to Presentation
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2601.07794v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly evaluated on their ability to perform multi-hop reasoning, i.e., to combine multiple pieces of...
The paper introduces a generation benchmark for culturally specific kinship terms, evaluating five open‑weight large language models (LLMs) on Hindi, Tamil, and Korean. Unlike prior multiple‑choice tests that treat kinship understanding as a recognition task, the study prompts LLMs to generate terms across two communicative tasks and compares results to a matched option‑supported baseline. Findings show that while models like GPT‑OSS120B and Llama‑3.370B can select correct terms in over 90% and 78% of cases respectively, they produce the correct term only 36% and 24% of the time, indicating a significant evaluation‑format gap and highlighting the difficulty of culturally specific kinship generation even when relationships are explicitly stated.
Current literature evaluates large language models (LLMs) on multilingual kinship understanding using multiple choice benchmarks, treating it as a recognition problem. We instead prompt five open weig...
arXiv:2608.28018v1 Announce Type: cross Abstract: Knowledge-intensive reasoning requires Large Language Models (LLMs) to ground answers in provided evidence. When evidence is insufficient, it is desi...
arXiv:2607. 23019v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting enables large language models (LLMs) to tackle multi-step reasoning tasks, yet the generated intermediate steps are not guaranteed to be logically sound.
arXiv:2604. 02512v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly exhibit human-like patterns of pragmatic and social reasoning.