Open research questions in Biomedical Text Mining and Ontologies
48 unresolved questions extracted from the limitations and future-work sections of 417 Biomedical Text Mining and Ontologies papers in our library. Each links back to the study that raised it.
What the literature leaves open
Abu-Salih B, AL-Qurishi M, Alweshah M, AL-Smadi M, Alfayez R, Saadeh H (2023) Healthcare knowledge graph construction: A systematic review of the state-of-the-art, open issues, and opportu- nities.
A Knowledge Graph Framework for Linking Health Assessment Scales and Scientific Literature in Low-Annotation Settings: Development and Evaluation Study · 2026 · DOIData availability statement The combined use of CiteSpace for structural network mapping and LDA for semantic modeling represents a methodologically complementary framework in bibliometric research. This dual ap- proach enables simultaneous characterization of intellectual linkages and thematic semantics, thereby strengthening interpretability. The observed correspondence between highly modular co-citation clus- ters and LDA-derived topics, particularly around “neoadjuvant immunotherapy” and “pathologic response”, supports the coherence of the field’s intellectual evolution. Interdisciplinary coupling among oncology, thoracic surgery, and immunology appears to have accel- erated translational progress and contributed to a globally connected research ecosystem. 4.8 Strengths, limitations, and future directions A key strength of this study is the integration of multiple bibliometric frameworks, enabling cross-validation between quanti- tative indicators (NP, TC, h-index) and relational metrics (total link strength, modularity, and coupling). Several limitations should be acknowledged. First, although we selected three major databases (WoS, PubMed, and Wanfang) to capture both English and Chinese literature, the exclusion of other important databases such as Scopus and Embase may introduce a risk of language and regional publication bias. Second, author disambiguation and citation lag can affect impact estimation, particularly for recently published trials and emerging themes. Third, bibliometric prominence primarily reflects patterns of academic attention and investment; high citation counts or keyword bursts should not be interpreted as evidence of clinical superiority of one regimen over another. Fourth, we did not perform a formal quality assessment or risk-of-bias evaluation on the included publications. All studies, regardless of their design, were assigned Publicly available datasets were analyzed in this study. This data can be found here: The data used and analyzed in this study are available from the corresponding author on reasonable request.
Global research trends and thematic evolution in neoadjuvant therapy and surgery for non-small cell lung cancer: a bibliometric and scientometric analysis · 2026 · DOIClinical NLP evaluation remains dominated by multiple-choice question answering (MCQA), which scores only final-answer accuracy and cannot detect when a model reaches the correct diagnosis while grounding it in irrelevant, absent, or contradictory evidence.
MIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical Exams · 2026Publication output was sparse before 2020, increased sharply during 2020-2021, and then stabilized at four to five publications annually during 2022-2024.
In the future, a hybrid model could be explored that combines general-purpose LLMs, which provide preliminary explanations, with domain-specific models that provide evidence links. The vast majority of studies utilize simulated or synthetic question- answering data, while studies employing real clinical text are scarce and primarily concentrated on predictive tasks.
Applications of natural language processing and large language models in sports injury assessment and rehabilitation decision-making: a scoping review · 2026 · DOIFuture work will focus on formal coverage evaluation, continued versioned governance as DS knowledge evolves, and assessment of whether this specialization workflow can be scaled to other rare diseases requiring integration of clinical, genetic, therapeutic, and translational knowledge.
Developing a Specialized Dravet Syndrome Ontology for Rare Disease Informatics and AI Applications · 2026 · DOIFuture development of RD-OMICS will focus on addressing current limitations and expanding its scope, depth and utility for RD research. One major priority is to incorporate richer contextual information into metadata harmonization and categorization. In particular, more systematic LLM-based approaches could leverage additional text sources, including experiment descriptions, sample processing details, library preparation protocols, and associated publications, to improve metadata interpretation, disease entity alignment, and assay classification. Incorporating this contextual reasoning would enable RD-OMICS to generate more accurate experiment-level metadata and better classify complex experimental designs, such as distinguishing bulk from single-cell sequencing studies. A second important direction is to expand the range of omics data sources beyond GEO. While GEO provides a valuable foundation for transcriptomics and other functional genomics datasets, bioRxiv preprint The copyright holder for this preprint (which was not certified by peer review) is the author/funder. This article is a US Government work. It is not subject to copyright under 17 USC 105 and is also made available for use under a CC0 license. https://doi.org/10.64898/2026.06.29.735296 this version posted July 3, 2026. doi:; RD omics data are distributed across multiple public repositories and institutional resources. Future versions of RD-OMICS will aim to integrate additional data sources such as the Sequence Read Archive, dbGaP, BioProject, ArrayExpress, PRIDE, MetaboLights, and other disease- or domain-specific repositories. Broadening the data source coverage will increase dataset completeness, reduce repository-specific bias, and enable more comprehensive discovery of RDrelevant omics studies across genomics, transcriptomics, epigenomics, proteomics, metabolomics, and single-cell modalities. Another key priority is to broaden disease coverage beyond the current set of 194 selected RDs. Future development will seek to include a larger and more diverse set of rare diseases by aligning disease mentions to standardized vocabularies and ontologies, such as MONDO, Orphanet, and OMIM. This expansion will improve representation of under-studied and underrepresented rare diseases, support cross-disease comparisons, and enable researchers to identify shared molecular mechanisms across phenotypically or genetically related conditions. Expanding disease coverage will also make RD-OMICS more useful for rare disease communities where data are sparse, fragmented, or difficult to discover.
Overall, while the use of AI poses numerous challenges, ethical recommendations for its application can be derived from the four core principles of psychological research ethics: respect for self-determination, non-maleficence, beneficence, and justice [67]. Concerns about privacy, informed consent, and adequate use of context fall under the principle of self- determination, while awareness of biases and data protection relate to non-maleficence. Transparency and copyright considerations are grounded in justice, requiring researchers to disclose their use of AI tools and ensure compliance with intellectual property norms. Finally, beneficence demands that AI serves a genuinely beneficial intent, and that its ecological costs are taken seriously. Training and running large AI models requires enormous amounts of energy and water, and depends on an infrastructure built on scarce physical resources [68–70]. These environmental costs are rarely visible to individual users, yet accumulate at a global scale, raising the question of whether the benefits of AI use in any given research context justify its ecological footprint. Ultimately, the use of AI should serve the well-being of all people, both now and in future generations. Hence, scientists should continue to monitor and carefully consider emerging ethical issues in this rapidly evolving landscape.
However, existing biomedical QA benchmarks mainly focus on exam-style knowledge, literature comprehension, or short-range multi-hop inference, leaving source-conditioned graph reasoning and evidence topology construction underexplored.
Weaving Multi-Source Evidence for Biomedical Reasoning: The BioMedHop Benchmark and BioWeave Framework · 2026The systematic review findings position domain-specialized NER from an exploratory phase to a consolidation phase, with transfer learning, fine-tuning, and knowledge integration becoming the prevailing paradigm, while ongoing challenges in semantic understanding and domain transferability present opportunities for future work on annotation-efficient methods, the formalization of domain characterization, and improving research coverage of underexplored application domains.
Systematic Review of Specialized NLP Pipelines for Domain-Specific Named Entity Recognition: Architectures, Efficacy, and Research Gaps · 2026 · DOIGenes/proteins and the three Gene Ontology categories aligned cleanly across PrimeKG and Hetionet (mutual coverage 94-99%), but disease overlap was sparse: only 0.
Beyond Identifier Matching: An Empirical Characterization of Failure Modes in Biomedical Knowledge Graph Integration · 2026 · DOIWe quantify how much concept overlap survives realistic alignment, and we characterize the new failure modes introduced by the methods that practitioners reach for when ID matching is insufficient.
Beyond Identifier Matching: An Empirical Characterization of Failure Modes in Biomedical Knowledge Graph Integration · 2026 · DOIWhile this study provides a comprehensive bibliometric mapping of the KAN research landscape, several method- ological limitations must be acknowledged to ensure the appropriate interpretation of its findings. First, the biblio- graphic data was retrieved exclusively from the Web of Sci- ence Core Collection database. Although WoS is widely recognized as one of the most authoritative sources for peer reviewed scientific literature, its exclusive use means that relevant publications indexed in other major databases, such as Scopus, IEEE Xplore, and Google Scholar, may not be captured in the present analysis. Studies published in conference proceedings or preprint repositories such as arXiv, which are particularly prevalent in the rapidly evolv- ing machine learning field, are also largely absent from the WoS Core Collection, potentially underrepresenting the full scope of KAN research activity. The restriction to Web of Science was a deliberate meth- odological decision to ensure that the bibliometric corpus consisted exclusively of peer reviewed publications with verified citation records, which is a standard and widely adopted criterion in bibliometric research. While this means that influential arXiv preprints and conference papers, including the foundational KAN paper by Liu et al. [25], are not part of the mapped corpus, their intellectual influ- ence is indirectly captured through the citation patterns of the journal articles that build upon them. Future studies may consider extending the corpus to include Scopus or IEEE Xplore to broaden coverage, particularly for conference- heavy subfields within the KAN literature. 1 3M. H. Sulaiman, Z. Mustaffa Second, the language restriction applied during data col- lection, which limited the dataset to English language pub- lications only, may have introduced a systematic bias by excluding potentially significant contributions published in other languages, particularly given the strong representation of research institutions from non-English speaking countries within the KAN literature. Third, the application of a mini- mum citation threshold of 9 for the bibliographic coupling analysis, while necessary to ensure network interpretability and analytical reliability, inevitably excludes a proportion of recently published contributions that have not yet accumu- lated sufficient citations despite their potential intellectual significance. This limitation is particularly relevant in a field as young as KAN research, where high quality publications from late 2025 and early 2026 may be systematically under- represented in the coupling network due to the inherent lag between publication and citation accumulation. Fourth, although the search string was carefully constructed to cap- ture both the theoretical and architectural dimensions of the field, the possibility remains that some relevant publications employing variant terminologies or unconventional key- word combinations were not retrieved, introducing a degree of search recall limitation that is difficult to fully quantify.
The Evolution of Kolmogorov-Arnold Networks: A Bibliometric Study of Theoretical Foundations and Engineering Applications · 2026 · DOIAbubakar Sadiq Muhammad, 2026, 14:3 ISSN (Online): 2348-4098 ISSN (Print): 2395-4752 International Journal of Science, Engineering and Technology An Open Access Journal Therefore, Furthermore, the study concludes that traditional neuroscience models, while foundational, are often limited by their inability to scale, adapt to heterogeneous data, or process unstructured text.
Future work may explore finer-grained knowledge injection mechanisms, more powerful multi-hop reasoning architectures, and stricter domain-adaptive training strategies to continually improve the system's performance in complex professional domains. Future work will explore multi-modal knowledge injection, dynamic knowledge base updates, and interpretable reasoning mechanisms to further enhance the system’s applicability, reliability, and explainability in real-world industrial scenarios, providing a feasible technical solution for intelligent question answering in the SEP domain and broader specialized vertical fields.
While our evaluation framework captures a broad spectrum of interpretive capabilities, several limitations should be acknowledged. First, the study focuses on retrospective task performance using structured prompts and predefined ground-truth corpora; it does not evaluate models in generative or forward-looking tasks such as hypothesis development, data harmonization, or longitudinal pattern inference. These use cases are critical to translational research and warrant dedicated methodological assessment. Second, although our semantic scoring pipeline incorporates embedding-based similarity and expert review, it remains constrained by the representation biases of the underlying language model used for vectorization. Semantically equivalent responses that diverge lexically may be underweighted if they fall below the cosine similarity threshold. Future approaches could incorporate multi-embedding ensembles or dynamic thresholding to improve sensitivity. Third, our study benchmarks general-purpose models accessed via public prompts. We did not include domain- specialized models fine-tuned on biomedical literature (e.g., PubMedBERT [22] or Galactica [23]) or retrieval-augmented generation (RAG) architectures, which may offer enhanced factual grounding and context integration. Including such sys- tems in future comparisons could provide valuable insight into the tradeoffs between generalizability and specialization. Finally, while our random baseline offers a stringent null model for statistical benchmarking, comparisons against estab- lished symbolic or rule-based information retrieval (IR) systems were not performed. Including classical baselines such as BM25 [24], concept graph traversal [25], or MeSH-based indexing could enrich future evaluations and help disentangle the contributions of semantic modeling versus domain priors [26].
La benchmarking large language models for extracting biobank-derived insights into health and disease · 2026 · DOIAs LLMs become increasingly embedded within biomedical research and clinical informatics pipelines, future evaluation paradigms must evolve to capture their full operational scope. Beyond entity retrieval and semantic matching, there is a pressing need to benchmark models on complex reasoning tasks, including multi-hop inference, causal explanation, and counterfactual generation [27]. Incorporating interpretability audits (such as attribution mapping and calibration diagnostics [28]) will also be critical for building trust and ensuring model transparency in high-stakes settings. Further, expanding evaluation to clinically actionable endpoints (such as treatment prioritization, risk score generation, or eligibility screening) would better align benchmarking efforts with translational objectives [29]. These tasks require not only semantic alignment but robust integration of structured medical logic, patient heterogeneity, and evolving clinical guidelines [30]. Finally, the incorporation of temporally-aware prompts, multimodal inputs (e.g., tabular biomarkers, imaging reports), and structured biomedical ontologies (e.g., SNOMED CT [31], UMLS, or MeSH) represents a promising frontier for enhancing factual grounding and context sensitivity [32]. Benchmarking models in these contexts will be essential for determining their readiness to support real-world biomedical decision-making, from cohort design to individualized care planning. The LLM landscape continues to evolve rapidly. The models evaluated here: Gemini 3 Pro, Claude Opus 4.5, Claude Sonnet 4, GPT-5.2, Mistral Large, and DeepSeek V3, represent the frontier as of January 2026. This evaluation updates our initial experiments conducted in late 2024 [33], which included earlier model generations (Gemini 2.0 Flash, Claude 3.5 Sonnet, GPT-4o, among others). The transition from our 2024 to 2026 evaluation demonstrates both the reusability of PLOS Computational Biology | https://doi.org/10.1371/journal.pcbi.1014224 April 20, 2026 15 / 18 our benchmark framework and the continued advancement in model capabilities. We encourage future studies to apply our open-source framework to evaluate subsequent model releases and contribute to longitudinal tracking of progress in biobank-related knowledge retrieval.
La benchmarking large language models for extracting biobank-derived insights into health and disease · 2026 · DOIKGBR development must incorporate methods to align mouse experimental evidence with human bone regeneration data, including standardized ontology mappings for orthologous genes, cell types, and signaling pathways across species to support translational research from preclinical to clinical applications.
Current validation of KGBR relies on manual expert spot checks and use-case plausibility assessments. Future work must establish quantitative benchmarks (e.g., link-prediction metrics, precision/recall for entity-relation pairs) evaluated against curated gold standards specific to bone regeneration domain concepts.
KGBR v0.5 is based solely on literature information and does not integrate data from electronic health records, radiomics, single-cell transcriptomics, or other multi-omics sources that would provide molecular-level mechanistic details and clinical phenotype data relevant to bone regeneration outcomes.
The current KGBR schema is SNOMED CT-centered and relatively shallow, lacking deeper integration of Gene Ontology and Disease Ontology. Additionally, extracted relations lack explicit temporal ordering, directionality constraints, and quantitative thresholds necessary to represent the dynamic nature of bone regeneration mechanisms.
Entity and relation extraction in KGBR may fail to capture complex material names and emerging cell states in bone regeneration, with extraction errors creating false links that are exacerbated at scale. The accuracy of entity-relation pairs extracted from biomedical literature needs quantitative evaluation against manually curated gold standards.
The current KGBR build relies exclusively on PubMed English abstracts and may miss older classic work, non-English publications, and interdisciplinary journals relevant to bone regeneration research. This language and database restriction creates gaps in coverage of seminal work and international research contributions to the field.
KGBR currently includes only literature published after 2020, excluding substantial foundational knowledge and canonical experimental evidence establishing BMP and Wnt signaling mechanisms that predate 2020. This temporal restriction causes the knowledge graph to overrepresent recent applications while underrepresenting the historical mechanistic foundation of bone regeneration signaling pathways.
This study investigated the use of state-of-the-art natural language processing and deep learning techniques for clinical named entity recognition from electronic health records and biomedical literature. Specifically, we compared two MCN-BERT models optimized with AdamP and AdamW against a BiLSTM model tuned with Hyperopt. Using two largebenchmark datasets, we aimed to automatically identify and classify medical entities like diseases, symptoms and adverse drug reactions from unstructured text. The experimental results demonstrate that the MCN-BERT approach optimized with AdamP achieved the best performance overall, attaining accura- cies of 99.58% and 96.15% on Datasets 1 and 2 respectively. The MCN-BERT model with AdamW optimization also delivered strong results, outperforming the BiLSTM baseline. Overall, our proposed domain-adapted trans- former architectures yielded superior clinical named entity recognition compared to prior work. These findings have important implications. By effectively extracting structured information from unstructured notes, clinical language models can support clinical decision making, drug safety surveillance, and knowledge discovery. Auto- matic identification of diseases and adverse events also paves the way for improved computational Phenotyping and pharmacovigilance. Looking ahead, further advances in model architectures and leveraging larger healthcare datasets hold promise to advance the state-of-the-art in medical natural language processing. The accuracy levels observed also suggest clinical language models are reaching maturity for real-world applications. Overall, our study underscores the growing potential of artificial intelligence to transform healthcare by unlocking insights from the tremendous amounts of textual patient data. This study demonstrated promising results for clinical named entity recognition using MCN-BERT models, however future research is still needed to advance these techniques for real-world clinical applications. Larger and more diverse healthcare datasets could be utilized to validate the generalizability of these models, particularly on underrepresented patient populations. Incorporating additional context like demographics, medical history and temporal trends can improve performance by provid- ing a more holistic view of the patient. Multi-task learning approaches that jointly solve related problems such as relationship extraction and coding assignment could generate a more comprehensive understanding compared Scientific Reports | (2024) 14:1507 | https://doi.org/10.1038/s41598-024-51615-5 21 Vol.:(0123456789)www.nature.com/scientificreports/ to single-task models. Leveraging self-supervised pre-training strategies has the potential to make better use of unlabeled clinical data. Integrating these models into clinical decision support systems and evaluating their impact on downstream tasks from diagnosis to treatment planning would help establish their clinical value. In the future work, we further plan to use another recent predictors such as pAtbP-EnC45, AIPs-SnTCN46, AFP- CMBPred47, cACP-DeepGram48, iACP-GAEnsC49, and Target-ensC_NP. Furthermore, we intended to used the CD-HIT tool was utilized to eliminate redundant peptide samples with homology50. Continued development of explainable AI is also important for gaining user trust in model-driven health- care. Further optimizing model architectures and expanding available data sources holds promise to consolidate medical language processing as a key enabling technology for advancing precision medicine through insights from patient narratives.
Most-cited papers in Biomedical Text Mining and Ontologies
- UniProt: the Universal Protein Knowledgebase in 2025 · Nucleic Acids Research · 2024 · 1,948 citations
- Mouse Genome Informatics: an integrated knowledgebase system for the laboratory mouse · Genetics · 2024 · 230 citations
- Comparative Toxicogenomics Database’s 20th anniversary: update 2025 · Nucleic Acids Research · 2024 · 188 citations
- PubTator 3.0: an AI-powered literature resource for unlocking biomedical knowledge · Nucleic Acids Research · 2024 · 141 citations
- OHDSI Standardized Vocabularies—a large-scale centralized reference ontology for international data harmonization · Journal of the American Medical Informatics Association · 2024 · 122 citations
- Optimizing classification of diseases through language model analysis of symptoms · Scientific Reports · 2024 · 114 citations
- Large Language Models Are Poor Medical Coders — Benchmarking of Medical Code Querying · NEJM AI · 2024 · 113 citations
- BioCLIP: A Vision Foundation Model for the Tree of Life · 2024 · 113 citations
- Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES): a method for populating knowledge bases using zero-shot learning · Bioinformatics · 2024 · 89 citations
- Differentiating ChatGPT-Generated and Human-Written Medical Texts: Quantitative Study · JMIR Medical Education · 2023 · 66 citations
Most recent work
- OmniCellAgent: An AI Scientist for Omic-Driven Scientific Discovery · bioRxiv · 2026
- Towards the explainability of protein language models · Nature Machine Intelligence · 2026
- Microbial Named Entity Recognition and Normalisation for AI-assisted Literature Review and Meta-Analysis · bioRxiv · 2026
- What level of automation is “good enough”? A benchmark of large language models for meta-analysis data extraction · Research Synthesis Methods · 2026
- High-precision Biomedical Text Corpora for Multi-Entity Recognition · bioRxiv · 2026
- Beyond the Gene in Genetics: How Isoform-Resolved Analysis Empowers the Study of Both Common and Rare Genetic Variation · medRxiv · 2026
- cadmus: a robust pipeline for scalable retrieval of full-text biomedical literature · bioRxiv · 2026
- scFAIR Consortium: a decentralized hub for single-cell RNA-Seq data standardization and unification · bioRxiv · 2026
- NMAstudio 2.0: An interactive tool for network meta-analysis to enhance understanding, interpretation, and communication of the findings · Research Synthesis Methods · 2026
- <scp>RescueGPT</scp> : An Automated System for Detecting Adverse Safety Events in Prehospital Emergency Medical Service Notes With a Zero‐Shot Approach With Large Language Models: A Proof‐of‐Concept Study · Learning Health Systems · 2026
Find a gap in your own Biomedical Text Mining and Ontologies sub-topic
This page shows what the Biomedical Text Mining and Ontologies literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →