arXiv Computation and Language

Enhancing Biomedical Named Entity Recognition via Multiple Programming Languages Instruction Tuning and Ensemble Method

The paper introduces MITE, a method that transforms biomedical named entity recognition (BioNER) into a structure‑to‑structure generation task by encoding instructions and outputs in multiple programming languages (Python, C++, Java). This approach provides structurally diverse supervision without extra biomedical knowledge, and during inference it aggregates predictions via entity‑level voting to reduce language‑specific variance. Experiments on six BioNER datasets show that MITE outperforms BERT‑based and LLM‑based baselines and generalizes well across datasets.

arXiv Computation and Language
4d ago

A Guideline-Augmented Multi-Agent Framework for Schema-as-Code Biomedical Named Entity Recognition

The paper introduces GAMA, a guideline-augmented multi-agent framework designed to improve biomedical named entity recognition (BioNER) using large language models (LLMs). GAMA constructs dataset-specific guideline memory by inducing and verifying annotation rules from training data, then employs a planning component to generate span-type hypotheses with rationales, a coding component to produce schema-constrained entity objects, and a verification module for structural compliance and dual-loop refinement. Experiments across five BioNER datasets demonstrate that GAMA consistently outperforms strong LLM-based baselines, with ablation studies confirming the effectiveness of each component.

By Songtao Li, Yijia Zhang, Shidi Zhang, Jianyuan Yuan, Fengyu Zhang, Hongfei Lin
arXiv Machine Learning
Sep 25

Language Specificity vs. Domain Diversity: Benchmarking Transformers for Bangla Medical NER

This study benchmarks transformer models for Bangla medical named entity recognition (NER), comparing BanglaBERT, multilingual BERT (mBERT), XLM‑RoBERTa, and GPT‑4o mini under zero‑shot and few‑shot prompting. Across a full test set of 3,179 samples, fine‑tuned XLM‑RoBERTa achieves a new state‑of‑the‑art F1‑score of 0.5959, while BanglaBERT lags with 0.4937, suggesting that domain diversity outweighs language specificity. The analysis shows high performance on Medicine and Specialist entities (F1 > 0.83) but lower accuracy on Symptoms (F1 0.4367), and demonstrates that fine‑tuned transformers outperform prompt‑only approaches by a factor of 3.76.

By Rakib Abdullah, Md. Maruful Islam Maruf
arXiv Machine Learning
Sep 10

Where Does the Signal Live? A Web Data Recipe for Medical Encoder Pretraining

The paper introduces a web‑data curation recipe for pretraining medical encoders, addressing the scarcity of large, diverse corpora in dense‑terminology domains like medicine. It proposes two complementary techniques: medical‑term density filtering to select documents rich in medical terminology, and signal‑amplifying rephrasing that uses an LLM to rewrite documents into denser variants with broader entity contexts. Applied to French medical NLP, the recipe produces the FineMed corpus and the DoctoBERT encoder family, achieving state‑of‑the‑art results on the DrBenchmark public benchmark and a proprietary clinical NER task.

By Bofeng Huang, Jacques Sun, Diane Bouchacourt, Nicolas Barascud, Fajwel Fogel
arXiv AI
Jul 13

TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

arXiv:2607. 08803v1 Announce Type: cross Abstract: The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology.

By Hyunjin Seo, Hyeon Hwang, Gyubok Lee, Jay Shin, Jimin Park, Taesoo Kim, Sanghoon Lee, Hongjoon Ahn, Sungjun Han, Sangwon Jung
arXiv AI
Sep 3

BioELX: Context-Aware Cross-lingual Biomedical Entity Linking without Task-Specific Supervision

BioELX is a retrieve‑rerank framework for cross‑lingual biomedical entity linking that tackles two key problems: the English‑biased UMLS alias training data and the degradation caused by naïvely adding context. It fine‑tunes SapBERT_multi with Wikidata‑derived cross‑lingual alias supervision to create shared concept neighborhoods, and then reranks candidates using pretrained LLMs with mention‑anchored prompting to focus on the target mention. Experiments demonstrate state‑of‑the‑art performance on four benchmarks, improving Recall@1 by 4.8–18.2 percentage points without task‑specific annotations.

By Yi Wang, Corina Dima, Liangyu Zhong, Steffen Staab
arXiv AI
Jun 11

Self-Prompting Small Language Models for Privacy-Sensitive Clinical Information Extraction

arXiv:2605. 04221v2 Announce Type: replace-cross Abstract: Clinical named entity recognition from dental progress notes is challenging because documentation is highly unstructured, domain-specific, and often privacy-sensitive.

By Yao-Shun Chuang, Tushti Mody, Uday Pratap Singh, Shirindokht Shiraz, Chun-Teh Lee, Ryan Brandon, Muhammad F Walji, Xiaoqian Jiang, Bunmi Tokede
arXiv AI
4d ago

SoftGene: Protein Language Model-Enhanced Soft Prompting for Interpretable Gene Set Annotation

SoftGene is a new framework that enhances gene set annotation by combining protein language model representations with a hybrid prompting scheme. It uses a hierarchical attention-based encoder built on ESM to encode gene sets from protein amino acid sequences, then merges soft prompts derived from these embeddings with hard prompts generated by a large language model. The approach is evaluated on Gene Ontology and MSigDB datasets, showing that integrating protein-sequence information with textual context improves overall annotation performance, though the benefit varies across biological domains.

By Drew Ross, Arya Hadizadeh Moghaddam, Dongjie Wang, Xiaoyu Zhang, Zijun Yao
arXiv Computation and Language
Sep 23

BELXTR: Biomedical Entity Linking via Contextualized Token Retrieval

BELXTR is a new biomedical entity linking model that uses a multi‑vector (late interaction) architecture to preserve token‑level matching information, unlike traditional embedding‑based approaches that compress mentions into a single vector. By extending the XTR model with a task‑specific training objective and active query expansion, BELXTR achieves state‑of‑the‑art performance on half of ten evaluated corpora, with an average 5‑percentage‑point gain in recall@1. The model shows especially strong results on cross‑species gene disambiguation, outperforming an LLM‑powered retrieve‑and‑rerank pipeline and approaching a specialized rule‑based system.

By Samuele Garda, Ulf Leser