arXiv AI By Mark Schutera

TINY_SCHILLER: A Drop-In German Drama Corpus for Small Language Models

Read the original on arXiv AI →

arXiv:2607. 19992v1 Announce Type: cross Abstract: tiny_schiller closes the small-language-model prototyping, fine-tuning, education, and research gap for German literary text, providing a single-file, drop-in counterpart to Karpathy's tiny_shakespeare.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
4d ago

SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models

SyntaxBench is a diagnostic benchmark and statistical evaluation framework for character‑level reasoning in large language models, comprising five core tasks—character counting, letter containment, palindrome detection, edit distance, and longest‑string selection—and a harder substring‑extraction stress test called index_to_span. The benchmark uses paired English and random‑string inputs, zero‑, one‑, and four‑shot prompts, and evaluates models from 2B to 32B parameters across multiple reasoning modes. It reports a wide range of metrics, including exact‑match and relaxed accuracy, Cohen’s kappa, McNemar tests, bootstrap confidence intervals, Kendall’s tau, class‑conditional metrics, tokenization analysis, and multiple‑comparison‑corrected tests.

By Mohsen Larni (Department of Computer Science, University of Nevada, Las Vegas), Sobhan Ebrahimi Azar (Department of Computer Science, University of Nevada, Las Vegas), Pouyan Nahed (Department of Computer Science, University of Nevada, Las Vegas), Kazem Taghva (Department of Computer Science, University of Nevada, Las Vegas)