arXiv AI By Jacob Nielsen, Stine L. Beltoft, Peter Schneider-Kamp, Lukas Galke Poech

SDUs DAISY: A Benchmark for Danish Culture

Read the original on arXiv AI →

SDUs DAISY is a factual knowledge benchmark focused on Danish cultural heritage, drawing topics from the Danish Culture Canon 2006. The dataset contains 741 manually verified closed‑ended question‑answer pairs generated by querying Wikipedia pages for each canon artifact, covering a wide historical span from 1300 BCE to contemporary pop music, design, and architecture. Baseline tests with state‑of‑the‑art language models show very low performance (best BLEU 0.17, F1 0.27), highlighting the benchmark’s difficulty and the need for improved cultural knowledge in AI systems.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 17

Knowledge-Graph Based Augmentation versus Retrieval Augmented Generation for Cultural-Related Question Answering

The paper compares Knowledge-Graph Based Augmentation (Graph-RAG) with Retrieval-Augmented Generation (RAG) for answering culturally specific questions. Using the LatamQA dataset, Graph-RAG, built automatically from Wikipedia via KGGen, matches RAG performance and reduces the base LLM’s error by 72% with a standard KG and 78% with a benchmark-aware variant. The approach also transfers zero‑shot to Portuguese, showing multilingual applicability.

By Pablo Poulenard, Yannis Karmim, Valentin Barri\`ere
arXiv Computer Vision
Aug 26

Wontopos Tablet 2: Measuring Multilingual and Multimodal Memory Retrieval Without Lexical Matching

The paper evaluates the tablet‑2 long‑term memory engine on multilingual text benchmarks and cross‑lingual photo retrieval without lexical matching. Tablet‑2 achieves high accuracy on LongMemEval‑S (95.7%) and moderate accuracy on BEAM‑1M (67.5%), with minimal variance across runs. In multimodal tests, it outperforms BM25 on image‑cell recall and shows significant language‑dependent performance gaps, especially for low‑resource languages.

By Sunwoo Kim