arXiv Computation and Language

EviScope: Paired Counterfactual Evidence Diagnostics for Faithful and Efficient Grounded Language Models

EviScope is a new paired counterfactual benchmark that evaluates grounded language models by fixing the question while manipulating evidence—adding, removing, distracting, or contradicting it. The v1.1 dataset includes 40 four‑condition quartets with repaired counterfactual claims and span‑level support labels for automated assessment. Experiments on Qwen2.5‑7B, Llama 3.1 8B, and Gemini 3.5 Flash show that paired metrics reveal grounding behaviors hidden by simple answer accuracy, such as unsupported answers, conflict blindness, and incorrect non‑answer actions.

arXiv AI
1d ago

AfriSyCo: Measuring Assertive Framing, Verification, and Wording Sensitivity Around African-Language Content

AfriSyCo investigates how different framing and verification strategies affect the accuracy of language models on African‑language factual content. The study uses a cross‑language factorial design with native‑language follow‑ups and English framing, analyzing 1,415 observations from 100 source questions across seven checkpoints and six languages. Results show that assertive framing boosts target selection by up to 30.4 points, while verification reduces it by 17.4 points, with strong interactions and large variability depending on wording and checkpoint.

By David Ababio Awuni, Rose-Mary Owusuaa Mensah Gyening, Elvis Gyasi Owusu