NIV: Neural Axis Variations for Variable Font Generation
arXiv:2606. 05261v1 Announce Type: cross Abstract: Variable fonts enable continuous variation of glyph geometry along semantic design axes such as weight, width, slant, and optical size.
arXiv:2606. 05261v1 Announce Type: cross Abstract: Variable fonts enable continuous variation of glyph geometry along semantic design axes such as weight, width, slant, and optical size.
LoGAN is a VLM-based agentic framework designed for few-shot multilingual font localization. It takes a handful of glyphs or logo letters and generates complete character sets across many languages, including CJK, by combining a glyph-level diffusion model, style finetuning, spacing/kerning transfer, and texture expansion. The method outperforms specialized font generators and state‑of‑the‑art image editors in glyph fidelity, style, texture, and kerning consistency on datasets covering more than 27 languages.
arXiv:2607. 20385v1 Announce Type: cross Abstract: Optical Character Recognition (OCR) for Persian remains substantially less mature than for Latin-script languages despite Persian being spoken by more than 110 million people across multiple countries.
arXiv:2609.21595v1 Announce Type: new Abstract: In-context learning using Large Language Models (LLMs) offers a compelling path to training-free post-OCR correction, yet its effectiveness for Devanag...
arXiv:2609.37141v1 Announce Type: new Abstract: Semantic typography is a design technique where the visual representation of a word conveys its semantic meaning, while maintaining its legibility. Exi...
The paper "When Explanations Cannot Be Read: Measuring and Correcting SHAP and LIME Rendering for Right-to-Left Languages" identifies that standard SHAP and LIME visualizations, designed for left‑to‑right scripts, fail to display attribution values correctly for right‑to‑left languages such as Urdu, Arabic, Persian, and Hebrew. It introduces SHAP‑RTL, a rendering layer that preserves the original attribution values while correcting reading direction, script shaping, and font selection for each language. The authors evaluate SHAP‑RTL on hate‑and‑offensive‑language datasets using TF‑IDF and logistic regression, showing that default rendering yields high character error rates, while SHAP‑RTL maintains correct visualizations across all tested languages.
The paper presents a local traditional OCR pipeline that can be iteratively fine‑tuned on both layout and appearance levels of complex historical Sanskrit manuscripts. By adapting to the specific manuscript distribution, the pipeline reduces human annotation effort and improves transcription accuracy across subsequent pages. The authors apply this method to three manuscripts, release a dataset with detailed layout and Unicode annotations in PAGE‑XML format, and benchmark the pipeline against leading multimodal large language models.
The paper presents a local traditional OCR pipeline that can be iteratively fine‑tuned on both layout and appearance levels of complex historical Sanskrit manuscripts. By adapting to the specific manuscript distribution, the pipeline improves transcription accuracy across subsequent pages, reducing the need for costly human annotation. The authors apply the method to three manuscripts, release a dataset with detailed layout and Unicode annotations in PAGE‑XML format, and benchmark the results against leading multimodal large language models.
arXiv:2607. 23319v1 Announce Type: cross Abstract: Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentations when applied to classical Indian languages.
Sinhala is a morphologically rich abugida spoken by roughly 16 million people in Sri Lanka, and to date, there are no publicly available real-world datasets for page-level Sinhala OCR. All previous studies for assessing Sinhala OCR models have used artificially generated data.
The paper shows that HuggingFace’s ByteLevel pre‑tokenizer, which treats a word as a sequence of Unicode letters, splits abugida scripts at every vowel sign, creating a training‑free lower bound on tokenizer fertility. Across 26 languages, all 17 abugidas exhibit increased token counts (up to 9×), while Latin, Cyrillic, Hangul, and Han remain unchanged. The authors demonstrate that correcting the character class reduces Nepali token counts, improves model performance, and that this issue is widespread in popular HuggingFace models.
IndicDetect is a benchmark for evaluating AI‑generated text detection in Hindi, Telugu, and Tamil. It pairs curated human‑written texts with LLM‑generated counterparts across multiple domains and generators, testing detectors under domain shift, generator shift, and adversarial perturbation. The study shows that supervised neural detectors fail to generalize to unseen generators and attacks, with Hindi experiencing the greatest degradation, indicating that robustness—not peak accuracy—is the main weakness in Indic language detectors.