arXiv AI

Instruction Finetuning DeepSeek-R1-8B Model Using LoRA and NEFTune

arXiv:2606. 10392v1 Announce Type: new Abstract: Financial named-entity recognition (NER) is essential for translating unstructured financial reports and news into structured knowledge graphs.

arXiv AI
Sep 17

A Scalable Framework for Automated NER Annotation Correction in Low-Resource Languages

The paper introduces a scalable, multi-step framework designed to improve the quality of Named Entity Recognition (NER) annotations, particularly in low-resource languages. It employs a frequency-based iterative approach that combines self‑training with a dual‑threshold mechanism to increase inference confidence. Experiments on various NER datasets show notable performance gains over the original data, and the study also investigates the use of generative Large Language Models for NER tasks.

By Toqeer Ehsan, Thamar Solorio
arXiv Computation and Language
Sep 1

Error-Type-Aware Loss Reweighting for Robust Named Entity Recognition with Noisy LLM Labels

Large language models (LLMs) are increasingly used to annotate datasets for training smaller, task‑specialized models such as named entity recognition (NER). However, current fine‑tuning processes ignore the annotation noise introduced by LLMs, leading to degraded performance, and existing noise‑robust losses fail to handle the heterogeneous nature of NER noise (e.g., missing mentions vs. type errors). The authors propose error‑type‑aware loss reweighting, which applies separate reweighting rules for different erroneous token types, improving F1 scores by 0.8–2.0 percentage points on average and up to 4.6 points on Wikigold at 24.1% noise.

By Elena Merdjanovska, Jonas Golde, Alan Akbik
arXiv Computation and Language
Sep 1

Generative vs. Encoder Models for Multilingual NER: A Comprehensive Empirical Study on Naamapadam

The paper compares generative and encoder-based neural models for multilingual Named Entity Recognition (NER) across the eleven languages of the Naamapadam benchmark. Five classic model families, four decoder-only large language models fine‑tuned with LoRA and 4‑bit NF4 quantisation, and nine generative models in zero‑to‑5‑shot inference were evaluated under strict CoNLL span‑level metrics. Encoder-based models (mBERT and XLM‑R) achieved substantially higher F1 scores—up to 0.675 on Hindi—than any generative architecture, with gaps of 7.5–40 percentage points; the best few‑shot result reached only 28% of the encoder baseline. The study identifies three language clusters (encoder‑dominant, partial‑coverage, and failure‑zone) and offers deployment guidelines based on transfer learning and low‑resource NLP principles.

By Jakkala Mahesh, Jatavath Shravan Kumar, Komalla Shivani, Sujoy Sarkar
arXiv Computation and Language
Sep 11

A Training-Free, Alignment-Free Approach to Corporate Intelligence: Application to SEC Filings

The paper introduces a training‑free, alignment‑free method for corporate intelligence that uses deterministic sparse seed vectors to hash word strings into a fixed high‑dimensional basis. By accumulating these seed vectors across sentence contexts, the authors create corpus‑specific semantic signatures that enable rapid document comparison, issuer fingerprinting, vocabulary shift tracking, and thematic sentence extraction—all on standard CPU hardware. Applied to a multi‑year set of SEC filings, the approach reveals distinct semantic profiles for major corporate events such as Boeing’s 737 MAX crisis, Intel’s supply‑chain disruptions, and Bunge’s acquisition of Viterra, with each profile traceable to its source sentences without any domain‑specific training or LLM inference.

By Jean-Fran\c{c}ois Delpech
arXiv Computation and Language
Aug 25

LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Data

LuxIT is a monolingual instruction‑tuning dataset for Luxembourgish, created by synthesizing instruction‑answer pairs from native texts using the DeepSeek‑R1‑0528 model and a quality‑assurance LLM‑as‑judge process. The resulting 227,507 high‑quality pairs were used to fine‑tune 14 LLMs (≤15 B parameters), yielding an average accuracy increase of +5.37 percentage points on standardized Luxembourgish proficiency exams and improvements in macro‑averaged F1 on nine of the fourteen downstream NLP tasks. These findings demonstrate that synthetic monolingual data can effectively enhance LLM performance in low‑resource languages and reveal the complex relationship between exam performance and practical NLP gains.

By Julian Valline, Cedric Lothritz, Siwen Guo, Jordi Cabot
arXiv Machine Learning
Aug 19

Where A Small Language Model Helps in Invoice Categorisation, Understood Through Embedding Geometry

The paper investigates using a small language model (SBERT) for invoice categorisation, showing that fine‑tuning on a single GPU yields 0.96 accuracy and 0.9 F1 with about 100 client‑specific invoices. It analyses the embedding geometry, finding that the sentence‑embedding space is globally anisotropic but locally isotropic, with clusters strongly linked to vendor identity. The study demonstrates that an in‑house SLM can outperform zero‑shot LLMs and vendor baselines while improving cost, security, and interpretability.

By Emma Ceccherini, Daniel Lawson, Anjulika Salhan
arXiv AI
Sep 15

Not all Negation Cues are Equal: Affixal Negations Yield Better Negation Understanding

The paper introduces NegCue, a large-scale dataset of 1.8 million samples that includes single-word, multi-word, and affixal negation cues, totaling over 200 unique forms. The authors pre-train encoder-only language models and large language models on this dataset to study how different negation types influence understanding. Experiments on five downstream benchmarks reveal that affixal negations provide the most significant performance gains, whereas single-word negations yield modest improvements, and that additional pre-training benefits both model types.

By Tian Tan, Eduardo Blanco
arXiv Machine Learning
Jun 2

GottBERT: a pure German Language Model

arXiv:2012. 02110v2 Announce Type: replace-cross Abstract: Pre-trained language models have significantly advanced natural language processing (NLP), especially with the introduction of BERT and its optimized version, RoBERTa.

By Raphael Scheible, Johann Frei, Fabian Thomczyk, Henry He, Patric Tippmann, Jochen Knaus, Victor Jaravine, Frank Kramer, Martin Boeker
Hugging Face Trending Papers
Aug 18

Where A Small Language Model Helps in Invoice Categorisation, Understood Through Embedding Geometry

The paper explores using a small language model (SBERT) for invoice categorisation, a task that requires nuanced accounting judgement. By analysing the embedding geometry of SBERT and DeBERTa, the authors find that the sentence‑embedding space is globally anisotropic but contains locally isotropic clusters tied to vendor identity. Fine‑tuned SBERT achieves 0.96 accuracy and 0.9 F1 with only about 100 client‑specific invoices, outperforming zero‑shot LLMs and vendor baselines, and demonstrates that in‑house SLMs can reduce cost, enhance security, and improve interpretability.