arXiv Machine Learning

Quantifying the Occult: A Comparative Study of Hindu and Buddhist Deities Using Machine Learning Methods

The paper presents a dual‑matrix computational framework that quantifies morphological and theological differences among 196 Hindu and Vajrayana Buddhist esoteric deities. It uses a Gower distance matrix with a new Cardinality Weighting algorithm for physical form and dense vector embeddings from LLMs for theological function, revealing how visual forms can mask shared functions and how high‑cardinality symbols cluster orthodox and Tantric entities. The study demonstrates near‑identical coordinates for the Hindu Chinnamasta and Buddhist Chinnamunda, and releases the architecture as an open‑source tool for Digital Humanities research.

Hugging Face Trending Papers
Aug 12

JieZi: A Large-Scale Expert-Audited Dataset and Benchmark for Ancient Chinese Character Exegesis

The scholarly exegesis of ancient Chinese characters demands integrating visual observation, linguistic analysis, and historical context. However, existing computational approaches focus narrowly on subtasks such as character recognition and retrieval, lacking the structured datasets and benchmarks required for comprehensive scholarly analysis.

arXiv Computation and Language
Aug 28

Which India Survives Translation? Narrative Homogenisation Across Indian Oral Traditions in LLMs

The study investigates how large language models (LLMs) handle diverse Indian oral traditions, using the Rajasthani Pabuji epic, Tamil Sangam poetry, and Bengali folk tales as case studies. By prompting Claude Sonnet and Gemini with 54 generation requests across generic, culturally specific, and regional-language prompts, the authors measured reference drift and cross-tradition convergence using Sentence‑BERT embeddings. Results show that while outputs stay closer to their own tradition than to others, there is significant cross‑tradition similarity (0.52–0.66), indicating partial homogenisation; moreover, regional‑language prompting consistently reduced fidelity to authentic traditions.

By Paarth Singh Rathore
Hugging Face Trending Papers
Jul 8

Billions of Sketches Reveal Hidden Cultural Variation in Human Concepts

Claims about the universality of human concepts have been predominantly assessed through linguistic similarity across languages and cultures. However, words are effective as communication devices because they compress rich experiential variation into shared conventions, potentially obscuring hidden individual and cultural differences in how concepts are mentally represented.

arXiv Computation and Language
Sep 3

MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models

MemeCULT-1K is a multilingual benchmark of 1,000 South Asian memes in Bengali, English, and Hindi, each paired with a cultural context note and three human-written explanations, plus an additional set of 54 Bengali regional dialect memes. The study evaluates thirteen vision‑language models under meme‑only and context‑aware settings, showing that providing minimal cultural context consistently improves performance across all models and languages. Error analysis indicates closed‑source models struggle with entity and reference misidentification, while open‑source models are limited by broader cultural knowledge gaps, especially in linguistic and phonological aspects.

By Tawsif Tashwar Dipto, Mehedi Ahamed, Radib Bin Kabir, Mueeze Al Mushabbir, Mohammed Saidul Islam, Mir Rayat Imtiaz Hossain, Md Tahmid Rahman Laskar, Sabbir Ahmed