Biochemistry, Genetics and Molecular Biology · Research topic

Open research questions in Machine Learning in Bioinformatics

133 unresolved questions extracted from the limitations and future-work sections of 373 Machine Learning in Bioinformatics papers in our library. Each links back to the study that raised it.

What the literature leaves open

  • The papers identify a critical shortage of high-quality, diverse, and well-annotated biological datasets as a barrier to training robust computational models, yet none proposes or evaluates a systematic approach to generating or curating such datasets at the scale and standardization level required by modern machine learning and AI methods in biology.

    Large language models in bioinformatics: a comprehensive survey · 2026 · DOI
  • Large language models and deep learning approaches are evaluated in this set only on sequence analysis, protein structure prediction, and drug design tasks; none of the papers applies these methods to the problem of inferring gene regulatory networks from multi-modal single-cell data, despite the availability of both the computational tools and the data types needed for such integration.

    Large language models in bioinformatics: a comprehensive survey · 2026 · DOI
  • However, the reliability of functional annotations from such searches has not been systematically quantified.

    Calibrated structural homology transfer yields putative molecular functions for domains of unknown function in four model proteomes · 2026 · DOI
  • Improvements in computational protein structure prediction have enabled searches for remote homologs of proteins whose molecular function remain unknown.

    Calibrated structural homology transfer yields putative molecular functions for domains of unknown function in four model proteomes · 2026 · DOI
  • However, the "black-box" nature of these representations limits transparency and explainability, posing challenges for human-AI collaboration and leaving open questions about their human-interpretable features.

    Sparse autoencoders uncover biologically interpretable features in protein language model representations · 2025 · DOI
  • Analysis of genomic and metagenomic sequences is inherently more challenging than that of amino acid sequences due to the higher divergence among evolutionarily related nucleotide sequences, variable k-mer and codon usage within and among genomes of diverse species, and poorly understood selective constraints.

    Enhancing nucleotide sequence representations in genomic analysis with contrastive optimization · 2025 · DOI
  • A key strength of Scorpio is its ability to generalize to novel DNA sequences and taxa, addressing a significant limitation of alignment-based methods.

    Enhancing nucleotide sequence representations in genomic analysis with contrastive optimization · 2025 · DOI
  • Previous studies were limited by the scarcity of data and the neglect of protein structures.

    Accurately predicting optimal conditions for microorganism proteins through geometric graph learning and language model · 2024 · DOI
  • The variable length of phage genome sequences. The need to incorporate positional information of k-mers. The comparison with existing state-of-the-art methods.

    PhageCGRNet: Integrating Chaos Game Representation of Genomes with Convolutional Neural Network for accurate phage host classification prediction · 2026 · DOI
  • To further evaluate the performance of PhageCGRNet on larger datasets. To explore the application of PhageCGRNet in other fields such as microbiology and genomics.

    PhageCGRNet: Integrating Chaos Game Representation of Genomes with Convolutional Neural Network for accurate phage host classification prediction · 2026 · DOI
  • The complexity of TRNs and the need for accurate reconstruction. The limited availability of annotated datasets for training and evaluation. The need for efficient and accurate methods for NER of TFs and regulated elements.

    Fine-tuned GPT-based foundation models effectively reconstruct bacterial transcriptional regulatory networks from literature · 2026 · DOI
  • Exploring the application of the approach to other bacteria and domains. Evaluating the performance of the approach using larger datasets and more complex TRNs.

    Fine-tuned GPT-based foundation models effectively reconstruct bacterial transcriptional regulatory networks from literature · 2026 · DOI
  • The gap in the current methods is the limitation of traditional experimental approaches for effector identification. The gap is also the need for more accurate and efficient methods for secreted effector prediction.

    Machine learning for the prediction of gram-negative bacterial secreted effectors: advances and challenges · 2026 · DOI
  • systems. Advances in uncertainty quantification will enable models to flag low-confidence predictions, guiding more efficient experimental validation. In conclusion, machine learning has become a powerful tool for large-scale effector discovery. By addressing current limitations and embracing emerging methodologies, researchers will be able to develop more accurate and robust prediction tools.

    Machine learning for the prediction of gram-negative bacterial secreted effectors: advances and challenges · 2026 · DOI
  • To consider the effect of various negative data sampling approaches on model performance. To exploit geometrical features and topological features through graph deep learning. To apply the proposed model to other related tasks.

    NeuroPred-GMC: a dual-branch deep learning architecture for neuropeptide prediction based on gated dilated convolutional network and multi-scale convolutional network · 2026 · DOI
  • The lack of efficient and intelligent computational methods for predicting neuropeptides. The limitations of conventional experimental methods for predicting neuropeptides.

    NeuroPred-GMC: a dual-branch deep learning architecture for neuropeptide prediction based on gated dilated convolutional network and multi-scale convolutional network · 2026 · DOI
  • The performance of UNKAI is limited for inputs that lack similar samples in the training set. The distribution of functional categories in the training set exhibits some bias. The model's predictive capabilities for entirely unseen EC subclasses are lower than the overall performance.

    UNKAI: A protein functional identity prediction model based on ESM-C latent representations and the attention mechanism · 2026 · DOI
  • To improve the performance of UNKAI for inputs that lack similar samples in the training set. To apply UNKAI to various fields, such as biochemistry, biophysics, and pharmacology. To develop new methods that can leverage pLMs and attention mechanisms for protein functional identity prediction.

    UNKAI: A protein functional identity prediction model based on ESM-C latent representations and the attention mechanism · 2026 · DOI
  • We do not claim parity with specialist remote-homology tools; published numbers for ESM-2, CATHe and PLMSearch on differently constructed splits reach 65--75%, and closing this gap is discussed as an open problem.

    OmniGene-4: A Unified Bio-Language MoE Model with Router-Level Interpretability · 2026 · DOI
  • The need for selective modulators to temper the activity of pathogenic T cells. The limited ability of traditional high-throughput screening to identify novel therapeutic agents.

    Few-shot learning-driven discovery of Lutein suppresses Th1-mediated inflammation via glucose metabolism · 2026 · DOI
  • The complexity of data preprocessing pipelines and the lack of unified implementations make methods and results difficult to reproduce and compare. The need for a unified framework to analyze different approaches for molecule retrieval from MS/MS spectra.

    MSAlign: Aligning Molecule and Mass Spectra Foundation Models for Metabolite Identification · 2026
  • The role of MERCs in carcinogenesis still remains unknown - The current interface should be interpreted as an informal pilot assessment rather than a formal usability study

    Cell Structure Segmentation in TEM Images of Murine Skin Melanoma Cells by Deep Learning Model · 2026 · DOI
  • Search for new tools facilitating the study of MERCs in tumor cells - Evaluation of the ability of the network models to differentiate cellular structures after pre-training on external datasets

    Cell Structure Segmentation in TEM Images of Murine Skin Melanoma Cells by Deep Learning Model · 2026 · DOI
  • The scarcity of clinical data limits the prediction of patient-level drug response. Biological differences introduce domain shifts that hinder direct translation to patient tumors.

    Deep learning for predicting patient drug response by transferring gene-level and cell-level knowledge to tumors · 2026 · DOI
  • Downstream applications of the dataset include comparative genomic and proteomic studies. The dataset can be used for training machine learning models to predict hormone activity, stability, or other bioactive properties.

    HORDB 2.0: a comprehensive dataset of peptide hormones · 2026 · DOI

Most-cited papers in Machine Learning in Bioinformatics

Most recent work

Find a gap in your own Machine Learning in Bioinformatics sub-topic

This page shows what the Machine Learning in Bioinformatics literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.

Open the Research Gap Finder →

Related topics in Biochemistry, Genetics and Molecular Biology

133 open questions have been extracted from the limitations and future-work passages of 373 Machine Learning in Bioinformatics papers in our library. Each one below links back to the study that raised it, so you can read the original claim in context.

Tools for your next paper

Compare the category — Honest roundups of the AI research tools, ours listed alongside the alternatives.

Command palette

Jump anywhere, run any action.