arXiv AI By Fabien Maury (Imagine - U1163, HeKA | U1346), Sol\`ene Grosdidier (Imagine - U1163), Maud de Dieuleveult (Imagine - U1163), Adrien Coulet (HeKA | U1346)

AAbAAC: An Annotated Corpus for Autoimmunity Information Extraction

Read the original on arXiv AI →

arXiv:2606. 13051v1 Announce Type: new Abstract: Despite advances in information extraction driven by deep learning and large language models, performance gaps remain in highly specialized biomedical fields, where domainspecific complexity poses challenges for generalist models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 2

MiNER: Fine-Tuned Biomedical Natural Language Processing for Malaria Disease Entity Recognition in Clinical Texts

The paper introduces MiNER, a fine‑tuned biomedical NLP system that uses BioBERT to extract malaria‑related named entities from scientific literature. It builds a large, annotated corpus of malaria articles, preprocesses the text, and applies supervised learning to improve extraction performance. Experiments show that MiNER outperforms other encoding and machine‑learning methods in precision, recall, and accuracy, and the authors release the human‑labeled dataset for further research.

By V. S. Anoop, Devika N
arXiv Computation and Language
Sep 22

Custom Named Entity Recognition and Topic Classification for Global Health Publications

This thesis explores how to select and adapt NLP models for global health literature when annotated data and computational resources are scarce. It compares skip‑gram word2vec models trained on increasingly large specialized corpora with BioWordVec for semantic tag discovery, finding that larger coverage does not always yield more useful domain associations. The study also evaluates convolutional spaCy models versus a RoBERTa transformer for named entity recognition, noting a trade‑off between higher F1 scores and longer inference time, and investigates MiniLM few‑shot versus BART‑MNLI zero‑shot classification for multi‑label topic classification, highlighting practical constraints of inference cost. "whyItMatters":"The work provides empirical guidance on balancing model accuracy and resource demands for building knowledge systems in low‑resource global health settings."

By Genis Skura, Antoine Geissb\"uhler, Jean-Luc Falcone
arXiv Machine Learning
Sep 25

Language Specificity vs. Domain Diversity: Benchmarking Transformers for Bangla Medical NER

This study benchmarks transformer models for Bangla medical named entity recognition (NER), comparing BanglaBERT, multilingual BERT (mBERT), XLM‑RoBERTa, and GPT‑4o mini under zero‑shot and few‑shot prompting. Across a full test set of 3,179 samples, fine‑tuned XLM‑RoBERTa achieves a new state‑of‑the‑art F1‑score of 0.5959, while BanglaBERT lags with 0.4937, suggesting that domain diversity outweighs language specificity. The analysis shows high performance on Medicine and Specialist entities (F1 > 0.83) but lower accuracy on Symptoms (F1 0.4367), and demonstrates that fine‑tuned transformers outperform prompt‑only approaches by a factor of 3.76.

By Rakib Abdullah, Md. Maruful Islam Maruf
arXiv AI
Jul 13

TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

arXiv:2607. 08803v1 Announce Type: cross Abstract: The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology.

By Hyunjin Seo, Hyeon Hwang, Gyubok Lee, Jay Shin, Jimin Park, Taesoo Kim, Sanghoon Lee, Hongjoon Ahn, Sungjun Han, Sangwon Jung