arXiv AI By Miguel Arana-Catania, Gillian Pink, Glenn Roe

Automatic Thematic Indexing of Large Literary Corpora: A Machine Learning Approach to Voltaire's Complete Works

Read the original on arXiv AI →

arXiv:2607. 09316v1 Announce Type: cross Abstract: Thematic indexing -- the practice of assigning structured conceptual labels to sections of text -- is essential to scholarly access in large-scale literary and historical editions, yet it remains a largely manual, labour-intensive process.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 18

Viveka-Insight: a cross-lingual concept graph and citation-grounded retrieval resource over the complete works of Swami Vivekananda in English and Bengali

Viveka-Insight is a bilingual resource and open‑source pipeline for Swami Vivekananda’s complete works, providing a structure‑preserving parse of 32,694 paragraphs and 168,842 sentences, a cross‑lingual concept graph with 8,362 language‑agnostic concepts, a bilingual alias inventory of 60,850 surface forms, and a human‑annotated set of 200 paragraph‑concept edges. The resource enables citation‑grounded retrieval across the English and Bengali corpora, achieving Recall@10 of 0.86 for known‑item cross‑lingual queries and demonstrating concept‑extraction precision of 0.60 (up to 0.71 with confidence filtering). The design is intended to be transferable to other multilingual classical corpora.

By Tamal Maharaj
arXiv Computation and Language
Sep 18

Improving Cross-Lingual Transfer for Sequential Sentence Classification in Research Papers via Structural Similarity

The paper investigates cross‑lingual transfer for sequential sentence classification (SSC) in research papers, focusing on 13 non‑English languages. Experiments show that linguistic proximity does not reliably predict transfer success, whereas structural similarity in rhetorical organization—particularly label distribution similarity—correlates positively with performance. The authors introduce three generative‑model methods that exploit structural cues, achieving parity with strong encoder baselines on‑domain and outperforming them when transferring to unseen languages.

By Kazuhiro Yamauchi, Marie Katsurai
arXiv Computation and Language
Sep 1

Loci Similes: A Benchmark for Extracting Intertextualities in Latin Literature

The paper introduces Loci Similes, a benchmark for detecting intertextual links in Latin literature. It provides a curated dataset of about 176,000 text segments and 1,490 expert-verified parallels, including 945 labeled references from an existing source. Baselines for retrieval and classification are established using both lexical methods and pretrained encoder language models.

By Julian Schelb, Michael Wittweiler, Marie Revellio, Barbara Feichtinger, Andreas Spitz
arXiv AI
Sep 2

Value Over Language Model: Detecting Original Contribution in Writing

The paper introduces VOLM, a framework that quantifies how much original value a human adds to a document beyond what a language model could generate from a task description alone. Unlike existing tools that focus on stylistic detection, VOLM extracts content at varying granularities, reconstructs it with an LLM, and compares these reconstructions to those derived from the task description. Evaluations across news articles, ICLR peer reviews, and argumentative essays show that VOLM can distinguish human-authored texts from LLM-generated ones while remaining robust to content-preserving transformations.

By Vibhhu Sharma, Thorsten Joachims, Sarah Dean
arXiv Machine Learning
Sep 17

TACTICS: Taxonomy-Aware Intelligent Corpus Sampling for Machine Translation

TACTICS is a method for selecting evaluation samples in machine translation that explicitly optimizes for coverage of rare linguistic categories, document-level coherence, and distributional fidelity to the full corpus. It builds a hierarchical taxonomy from a locale style guide, classifies segments, and chooses a fixed-budget subset that better represents the full range of phenomena a system must handle. Compared to random, lexical, or embedding-based selection, TACTICS improves coverage of rare categories and yields more accurate system rankings with fewer segments.

By Prasanth Bathala, Anubhav Shrimal, Sukhdeep Singh Kharbhanda, Pradyumna Lanka, Rohit Dhaipule