EconPapers    
Economics at your fingertips  
 

CellExLink: End-to-end cell-type recognition and normalization in biomedical text

Alimire Nabijiang and Leili Shahriyari

PLOS Computational Biology, 2026, vol. 22, issue 7, 1-13

Abstract: Cell types are described in biomedical literature using diverse names, abbreviations, and phenotype phrases, which complicates their recognition and normalization. We developed CellExLink, an end-to-end pipeline that identifies cell-type mentions and normalizes them to Cell Ontology (CL) identifiers. The recognizer was fine-tuned and evaluated on five heterogeneous biomedical corpora spanning full-length articles, article excerpts, figure captions, abstracts, and anatomical text passages. These resources include fine-grained phenotype-defined populations, heterogeneous cell populations, and abbreviated mentions. Across the five corpora, CellExLink achieved macro-average exact- and relaxed-span F1 scores of 0.766 and 0.855, respectively. For CL identifier normalization on gold-standard mention spans, F1 scores ranged from 0.690 to 0.874. In strict end-to-end evaluation, which required both an exact mention span and the correct CL identifier, F1 scores ranged from 0.552 on a figure-caption corpus to 0.725 on a corpus of full-text article excerpts. CellExLink outperformed the evaluated off-the-shelf systems in cell mention recognition, CL identifier normalization, and end-to-end extraction. By converting unannotated biomedical text into cell-type spans linked to standardized CL identifiers, CellExLink provides a practical foundation for downstream applications, including literature curation, relation extraction, and knowledge graph construction.Author summary: Cell types are central to biomedical research, but biomedical papers often use different names, abbreviations, and synonyms for the same cell type. This variation makes it difficult for automated processes to collect and compare cell-type information across papers. Reliable automated extraction is important because literature mining requires consistent cell-type identification before evidence from different studies can be searched, integrated, or reused. Existing off-the-shelf biomedical text-mining tools provide useful functionality, but their ability to support cell-type extraction remains limited and inconsistent. To address this gap, we developed CellExLink, a pipeline that finds cell-type entities in biomedical text and links them to standard Cell Ontology identifiers. We evaluated the pipeline on several biomedical corpora and compared it with existing tools that support cell-type extraction. Across these evaluations, CellExLink showed clear accuracy gains in both detecting cell-type entities and assigning correct standard identifiers. Together, these gains make CellExLink useful for extracting more reliable standardized cell-type information from large collections of papers, supporting literature curation, relation extraction, knowledge graph construction, and studies of cell-type-specific roles in diseases, drug responses, and biological pathways.

Date: 2026
References: Add references at CitEc
Citations:

Downloads: (external link)
https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1014556 (text/html)
https://journals.plos.org/ploscompbiol/article/fil ... 14556&type=printable (application/pdf)

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:plo:pcbi00:1014556

DOI: 10.1371/journal.pcbi.1014556

Access Statistics for this article

More articles in PLOS Computational Biology from Public Library of Science
Bibliographic data for series maintained by ploscompbiol ().

 
Page updated 2026-07-26
Handle: RePEc:plo:pcbi00:1014556