arXiv AI

Domain-Specific Text Embedding Models for Entity Resolution

arXiv:2608. 16161v1 Announce Type: cross Abstract: General-purpose text embedding models are designed to capture semantic similarity but are not optimised for distinguishing entity records that represent the same real-world business or person.

arXiv Computation and Language
Sep 7

LentEx: Generalizable Latent Entity Extraction via Synthetic Data and Instruction-Tuned LLMs

LentEx is a new framework for latent entity extraction that uses synthetic data generation and instruction fine‑tuning to train smaller, efficient large language models. By creating diverse, contextually rich synthetic examples through a template‑based approach, LentEx overcomes the lack of labeled datasets and achieves strong performance, surpassing state‑of‑the‑art models on the MTEB Clustering Benchmark. The method also generalizes well to unseen domains, making it useful for tasks such as retrieval‑augmented generation, customer persona analysis, and knowledge graph enrichment.

By Umesh Bodhwani, Yuan Ling, Cibi Chakravarthy Senthilkumar, Shujing Dong, Yarong Feng, Hongfei Li, Ayush Goyal
arXiv Computation and Language
Sep 23

BELXTR: Biomedical Entity Linking via Contextualized Token Retrieval

BELXTR is a new biomedical entity linking model that uses a multi‑vector (late interaction) architecture to preserve token‑level matching information, unlike traditional embedding‑based approaches that compress mentions into a single vector. By extending the XTR model with a task‑specific training objective and active query expansion, BELXTR achieves state‑of‑the‑art performance on half of ten evaluated corpora, with an average 5‑percentage‑point gain in recall@1. The model shows especially strong results on cross‑species gene disambiguation, outperforming an LLM‑powered retrieve‑and‑rerank pipeline and approaching a specialized rule‑based system.

By Samuele Garda, Ulf Leser
arXiv Machine Learning
Sep 7

Nepali Passport Question Answering: A Low-Resource Dataset for Public Service Applications

The paper introduces a Nepali Question‑Answer dataset focused on passport‑related FAQs to support information retrieval in a low‑resource language. The authors fine‑tune transformer‑based embedding models for semantic similarity and compare them against the BM25 baseline. Their experiments show that fine‑tuned SBERT models outperform BM25, while multilingual E5 embeddings achieve the best overall retrieval performance.

By Funghang Limbu Begha, Praveen Acharya, Bal Krishna Bal
arXiv Computation and Language
Sep 22

ChemCLIR-Bench: Benchmarking Cross-Lingual Information Retrieval in Multilingual Chemical Patents

arXiv:2609.23231v1 Announce Type: new Abstract: Cross-lingual information retrieval (CLIR) is increasingly important in multi-national industries, where critical technical evidence may exist in a dif...

By Mahdi Astaraki, Mohammad Khodadad, Reza Namazi, Mohammad Arshi Saloot, Amir Reza Behzad Moghadam, Hamidreza Mahyar, Soheila Samiee
arXiv AI
Sep 10

Beyond Single-Negative Preference: Multi-Negative DPO for LLM-Centric Historical Entity Linking

The paper introduces Multi-Negative Direct Preference Optimisation (MDPO), a pairwise objective that compares the correct entity with all valid rejected candidates for each mention, extending the single-negative approach used in prior work. MDPO retains the Bradley‑Terry formulation of Direct Preference Optimisation while leveraging the full candidate set through masked, length‑normalised sequence scores. Experiments on French, German, English, Swedish, and Finnish historical newspaper datasets (hipe‑2020 and newseye) show that MDPO outperforms both supervised fine‑tuning and single‑negative DPO, especially for NIL mentions, semantic ambiguity, OCR noise, and historically challenging names, and highlight candidate retrieval as a key bottleneck.

By Tien Nam Nguyen, Emanuela Boros, Ahmed Hamdi, Adam Jatowt, Micka\"el Coustaty, Antoine Doucet