Open research questions in Machine Learning in Bioinformatics
133 unresolved questions extracted from the limitations and future-work sections of 373 Machine Learning in Bioinformatics papers in our library. Each links back to the study that raised it.
What the literature leaves open
The papers identify a critical shortage of high-quality, diverse, and well-annotated biological datasets as a barrier to training robust computational models, yet none proposes or evaluates a systematic approach to generating or curating such datasets at the scale and standardization level required by modern machine learning and AI methods in biology.
Large language models and deep learning approaches are evaluated in this set only on sequence analysis, protein structure prediction, and drug design tasks; none of the papers applies these methods to the problem of inferring gene regulatory networks from multi-modal single-cell data, despite the availability of both the computational tools and the data types needed for such integration.
However, the reliability of functional annotations from such searches has not been systematically quantified.
Calibrated structural homology transfer yields putative molecular functions for domains of unknown function in four model proteomes · 2026 · DOIImprovements in computational protein structure prediction have enabled searches for remote homologs of proteins whose molecular function remain unknown.
Calibrated structural homology transfer yields putative molecular functions for domains of unknown function in four model proteomes · 2026 · DOIHowever, the "black-box" nature of these representations limits transparency and explainability, posing challenges for human-AI collaboration and leaving open questions about their human-interpretable features.
Sparse autoencoders uncover biologically interpretable features in protein language model representations · 2025 · DOIAnalysis of genomic and metagenomic sequences is inherently more challenging than that of amino acid sequences due to the higher divergence among evolutionarily related nucleotide sequences, variable k-mer and codon usage within and among genomes of diverse species, and poorly understood selective constraints.
Enhancing nucleotide sequence representations in genomic analysis with contrastive optimization · 2025 · DOIA key strength of Scorpio is its ability to generalize to novel DNA sequences and taxa, addressing a significant limitation of alignment-based methods.
Enhancing nucleotide sequence representations in genomic analysis with contrastive optimization · 2025 · DOIPrevious studies were limited by the scarcity of data and the neglect of protein structures.
Accurately predicting optimal conditions for microorganism proteins through geometric graph learning and language model · 2024 · DOIThe variable length of phage genome sequences. The need to incorporate positional information of k-mers. The comparison with existing state-of-the-art methods.
PhageCGRNet: Integrating Chaos Game Representation of Genomes with Convolutional Neural Network for accurate phage host classification prediction · 2026 · DOITo further evaluate the performance of PhageCGRNet on larger datasets. To explore the application of PhageCGRNet in other fields such as microbiology and genomics.
PhageCGRNet: Integrating Chaos Game Representation of Genomes with Convolutional Neural Network for accurate phage host classification prediction · 2026 · DOIThe complexity of TRNs and the need for accurate reconstruction. The limited availability of annotated datasets for training and evaluation. The need for efficient and accurate methods for NER of TFs and regulated elements.
Fine-tuned GPT-based foundation models effectively reconstruct bacterial transcriptional regulatory networks from literature · 2026 · DOIExploring the application of the approach to other bacteria and domains. Evaluating the performance of the approach using larger datasets and more complex TRNs.
Fine-tuned GPT-based foundation models effectively reconstruct bacterial transcriptional regulatory networks from literature · 2026 · DOIThe gap in the current methods is the limitation of traditional experimental approaches for effector identification. The gap is also the need for more accurate and efficient methods for secreted effector prediction.
Machine learning for the prediction of gram-negative bacterial secreted effectors: advances and challenges · 2026 · DOIsystems. Advances in uncertainty quantification will enable models to flag low-confidence predictions, guiding more efficient experimental validation. In conclusion, machine learning has become a powerful tool for large-scale effector discovery. By addressing current limitations and embracing emerging methodologies, researchers will be able to develop more accurate and robust prediction tools.
Machine learning for the prediction of gram-negative bacterial secreted effectors: advances and challenges · 2026 · DOITo consider the effect of various negative data sampling approaches on model performance. To exploit geometrical features and topological features through graph deep learning. To apply the proposed model to other related tasks.
NeuroPred-GMC: a dual-branch deep learning architecture for neuropeptide prediction based on gated dilated convolutional network and multi-scale convolutional network · 2026 · DOIThe lack of efficient and intelligent computational methods for predicting neuropeptides. The limitations of conventional experimental methods for predicting neuropeptides.
NeuroPred-GMC: a dual-branch deep learning architecture for neuropeptide prediction based on gated dilated convolutional network and multi-scale convolutional network · 2026 · DOIThe performance of UNKAI is limited for inputs that lack similar samples in the training set. The distribution of functional categories in the training set exhibits some bias. The model's predictive capabilities for entirely unseen EC subclasses are lower than the overall performance.
UNKAI: A protein functional identity prediction model based on ESM-C latent representations and the attention mechanism · 2026 · DOITo improve the performance of UNKAI for inputs that lack similar samples in the training set. To apply UNKAI to various fields, such as biochemistry, biophysics, and pharmacology. To develop new methods that can leverage pLMs and attention mechanisms for protein functional identity prediction.
UNKAI: A protein functional identity prediction model based on ESM-C latent representations and the attention mechanism · 2026 · DOIWe do not claim parity with specialist remote-homology tools; published numbers for ESM-2, CATHe and PLMSearch on differently constructed splits reach 65--75%, and closing this gap is discussed as an open problem.
The need for selective modulators to temper the activity of pathogenic T cells. The limited ability of traditional high-throughput screening to identify novel therapeutic agents.
Few-shot learning-driven discovery of Lutein suppresses Th1-mediated inflammation via glucose metabolism · 2026 · DOIThe complexity of data preprocessing pipelines and the lack of unified implementations make methods and results difficult to reproduce and compare. The need for a unified framework to analyze different approaches for molecule retrieval from MS/MS spectra.
MSAlign: Aligning Molecule and Mass Spectra Foundation Models for Metabolite Identification · 2026The role of MERCs in carcinogenesis still remains unknown - The current interface should be interpreted as an informal pilot assessment rather than a formal usability study
Cell Structure Segmentation in TEM Images of Murine Skin Melanoma Cells by Deep Learning Model · 2026 · DOISearch for new tools facilitating the study of MERCs in tumor cells - Evaluation of the ability of the network models to differentiate cellular structures after pre-training on external datasets
Cell Structure Segmentation in TEM Images of Murine Skin Melanoma Cells by Deep Learning Model · 2026 · DOIThe scarcity of clinical data limits the prediction of patient-level drug response. Biological differences introduce domain shifts that hinder direct translation to patient tumors.
Deep learning for predicting patient drug response by transferring gene-level and cell-level knowledge to tumors · 2026 · DOIDownstream applications of the dataset include comparative genomic and proteomic studies. The dataset can be used for training machine learning models to predict hormone activity, stability, or other bioactive properties.
Most-cited papers in Machine Learning in Bioinformatics
- Evolutionary-scale prediction of atomic-level protein structure with a language model · Science · 2023 · 4,885 citations
- A novel genetic system to detect protein–protein interactions · Nature · 1989 · 4,762 citations
- Improved protein structure prediction using potentials from deep learning · Nature · 2020 · 3,127 citations
- Highly accurate protein structure prediction for the human proteome · Nature · 2021 · 2,936 citations
- SignalP 6.0 predicts all five types of signal peptides using protein language models · Nature Biotechnology · 2022 · 2,929 citations
- Large language models generate functional protein sequences across diverse families · Nature Biotechnology · 2023 · 1,095 citations
- iPHoP: An integrated machine learning framework to maximize host prediction for metagenome-derived viruses of archaea and bacteria · PLoS Biology · 2023 · 492 citations
- Enzyme function prediction using contrastive learning · Science · 2023 · 421 citations
- Fast, accurate antibody structure prediction from deep learning on massive set of natural antibodies · Nature Communications · 2023 · 304 citations
- SBSM-Pro: support bio-sequence machine for proteins · Science China Information Sciences · 2024 · 272 citations
Most recent work
- Protein and genomic language models uncover the unexplored diversity of bacterial immunity · Science · 2026
- The Dayhoff Atlas: scaling sequence diversity for improved protein generation · bioRxiv · 2026
- nail: software for high-speed sequence annotation with profile hidden Markov models · bioRxiv · 2026
- Understanding language model scaling for protein fitness prediction · Nature Computational Science · 2026
- Beyond additivity: zero-shot methods cannot predict impact of epistasis on protein properties and function · bioRxiv · 2026
- Rewriting protein alphabets with language models · bioRxiv · 2026
- Advances in protein function prediction from the fifth CAFA challenge · bioRxiv · 2026
- Emergence of Biological Structural Discovery in General-Purpose Language Models · bioRxiv · 2026
- A hybrid model for multiscale prediction and generation of phase separating protein regions · bioRxiv · 2026
- Biocentral: Embedding-based Protein Predictions · Journal of Molecular Biology · 2026
Find a gap in your own Machine Learning in Bioinformatics sub-topic
This page shows what the Machine Learning in Bioinformatics literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →