Biochemistry, Genetics and Molecular Biology · Research topic

Open research questions in Machine Learning in Bioinformatics

37 unresolved questions extracted from the limitations and future-work sections of 268 Machine Learning in Bioinformatics papers in our library. Each links back to the study that raised it.

What the literature leaves open

  • <title>Abstract</title> Predicting protein-ligand binding affinity from three-dimensional (3D) structural data is a central task in structure-based drug discovery, yet it remains challenging due to limited data availability, structural complexity, and the sparse nature of 3D molecular representations.

    Application of vision transformers to protein-ligand affinity prediction · 2026 · DOI
  • The dodecin signal is therefore a partial-subdomain match limited by the size of the dodecin chain rather than a full-topology alternative. The heatmap is sparse, with most cells exactly zero, and the non-zero signal concentrates on the top rows of the plot in the same four cases the prose above describes, namely P.

    Tokens, Topologies, Taxa: Towards Declarative Biology and Bioengineering · 2026 · DOI
  • Experimental determination of all possible mutation effects is infeasible, and while state-of-the-art tools such as AlphaMissense show promise, their diagnostic performance is insufficient and they are often difficult to run locally.

    pLM-SAV: A Δ-Embedding Approach for Predicting Pathogenic Single Amino Acid Variants · 2026 · DOI
  • To address the limited sequence diversity, sparse biological grounding, and the still poorly understood mechanisms underlying CPPs uptake, we constructed CPP2Vec-GenSet, a hybrid dataset that integrates computationally generated peptides with experimentally curated CPPs.

    CPP2Vec: A representation learning approach for cell-penetrating peptides prediction · 2026 · DOI
  • Weight-locking with spectral deformation has been proposed as a potential method to prevent fine-tuning of neural networks, but has not been systematically evaluated in biological AI models.

    Safeguarding open-weight genomic foundation models through weight locking · 2026 · DOI
  • By operating solely on genomic sequence, DDTRN provides a scalable, interpretable, and data-efficient framework for bacterial TRN inference in species where expression data are scarce, and it establishes a foundation for future multimodal integration with condition-specific regulatory information.

    DDTRN: Predicting Bacterial Transcriptional Regulatory Networks Based on Gene Sequences using Dual Descriptor · 2026 · DOI
  • Although the OGTs for most bacteria remain unknown, the increasing availability of genomes from uncultivated and cultivated taxa has made it advantageous to build genomic, cultivation-independent models to infer OGT.

    Predicting optimal growth temperatures of bacteria using learned structural information from a single protein · 2026 · DOI
  • Significance StatementMutational effects prediction with protein language models tends to vary widely in prediction accuracy, depending on the dataset considered.

    Intrinsic dataset features drive mutational effect prediction by protein language models · 2026 · DOI
  • While protein generation has broadly followed trends in NLP, two directions remain underexplored: alignment methods that optimize model behavior toward design objectives, and prompting-based control at inference time without fine-tuning.

    ProtGPT3: an Open-source family of Promptable and Aligned Protein Language Models · 2026 · DOI
  • We do not claim parity with specialist remote-homology tools; published numbers for ESM-2, CATHe and PLMSearch on differently constructed splits reach 65--75%, and closing this gap is discussed as an open problem.

    OmniGene-4: A Unified Bio-Language MoE Model with Router-Level Interpretability · 2026 · DOI
  • systems. Advances in uncertainty quantification will enable models to flag low-confidence predictions, guiding more efficient experimental validation. In conclusion, machine learning has become a powerful tool for large-scale effector discovery. By addressing current limitations and embracing emerging methodologies, researchers will be able to develop more accurate and robust prediction tools.

    Machine learning for the prediction of gram-negative bacterial secreted effectors: advances and challenges · 2026 · DOI
  • for their translocation and biological ML has been widely applied to address key challenges in secreted effector identification and characterization. Numerous candidate effectors discovered using ML models have been experimentally functions, validated demonstrating the value of ML in effector discovery. ML will likely continue to accelerate research in this field, helping scientists elucidate pathogenic mechanisms and prioritize candidates for antimicrobial therapies and vaccine development. Based on the limitations discussed above, we outline three directions for future work: (i) multimodal dataset construction, (ii) meta- learning to address data scarcity and class imbalance, and (iii) generalizability assessment and uncertainty quantification.

    Machine learning for the prediction of gram-negative bacterial secreted effectors: advances and challenges · 2026 · DOI
  • Database resources The accuracy of prediction models relies heavily on the quality and coverage of training datasets. Over the past 2 decades, multiple experimentally validated resources have been curated to support ML-based effector prediction. These resources fall into two focus on categories, namely, system-specific databases individual secretion pathways and cross-system databases that integrate data across multiple secretion systems. A summary of commonly used databases is provided in Table 1.

    Machine learning for the prediction of gram-negative bacterial secreted effectors: advances and challenges · 2026 · DOI
  • The fine-tuned GPT model requires pre-defined output categories for regulatory effects during training and inference. The framework cannot discover or predict novel or context-specific regulatory mechanisms not represented in the training data, limiting its application to TRN reconstruction when organisms employ unconventional or understudied regulatory effects beyond activation, repression, and the unrecognized derepression category.

    Fine-tuned GPT-based foundation models effectively reconstruct bacterial transcriptional regulatory networks from literature · 2026 · DOI
  • The study manually curated only 250 false positive sentences representing 56 unique regulatory interactions to characterize model errors. A larger-scale systematic curation effort with stratified sampling across the full false positive dataset is needed to quantify the distribution of the nine error categories and determine which patterns (e.g., gene-product encoding, derepression) require targeted model improvements in the GPT-based TRN reconstruction pipeline.

    Fine-tuned GPT-based foundation models effectively reconstruct bacterial transcriptional regulatory networks from literature · 2026 · DOI
  • While the authors propose applying the BERN2-based algorithm with species-specific transcription factor and regulated element dictionaries to reconstruct TRNs in diverse bacterial species, no evaluation has been performed on non-Salmonella bacteria. The scalability and accuracy of the fine-tuned GPT approach across clinically and biologically relevant bacteria with different regulatory mechanisms, genetic organization, and literature coverage patterns requires empirical validation.

    Fine-tuned GPT-based foundation models effectively reconstruct bacterial transcriptional regulatory networks from literature · 2026 · DOI
  • The fine-tuned GPT model does not recognize the 'derepressor' regulatory effect category, which accounts for 14% of false positives in the Salmonella TRN reconstruction. The model must be retrained or augmented to explicitly handle derepression mechanisms where transcription factors directly bind and remove repression from target promoters, a biologically distinct category from simple repression.

    Fine-tuned GPT-based foundation models effectively reconstruct bacterial transcriptional regulatory networks from literature · 2026 · DOI
  • The GPT-based model achieves only 30% true positive rate among false positive predictions in manual curation of 250 sentences from the Salmonella TRN reconstruction. Future work must systematically address the 52% category of false positives where no regulatory interaction is expressed between transcription factors and regulated elements, particularly sentences describing gene-to-protein relationships that are incorrectly classified as activation or repression events.

    Fine-tuned GPT-based foundation models effectively reconstruct bacterial transcriptional regulatory networks from literature · 2026 · DOI
  • The contrastive learning framework empowers CLEAN to confidently (i) annotate understudied enzymes, (ii) correct mislabeled enzymes, and (iii) identify promiscuous enzymes with two or more EC numbers-functions that we demonstrate by systematic in silico and in vitro experiments.

    Enzyme function prediction using contrastive learning · 2023 · DOI
  • Protein language models have been increasingly successful on tasks ranging from fitness prediction to functional design, yet what biological knowledge they acquire and where it is encoded within their internal representations remain underexplored.

    High-resolution dissection of concept acquisition in different families of protein language models · 2026 · DOI
  • However, the complexity and lack of standardized data preparation and bioinformatic analyses present significant challenges, especially for newcomers to the domain.

    ALPAR: Automated Learning Pipeline for Antimicrobial Resistance · 2026 · DOI
  • Genomic foundation models such as Evo 2 are increasingly applied to microbial genomics, yet how well their representations capture viral genome organisation, and how reliably they generate viral sequence, remain poorly characterised.

    Genomic foundation model embeddings encode higher-order viral genome architecture beyond sequence composition: a benchmark of Evo 2 · 2026 · DOI
  • 89) reveals a key limitation of the current framework, which may be overcome through supervised training.

    Predicting Protein Electrostatics with Protein Language Models · 2026 · DOI
  • Data augmentation is a promising remedy, but systematic evaluation of augmentation strategies for protein sequences remains limited, and the conditions under which augmentation confers downstream benefits are not well understood.

    ProtAug: An Empirical Investigation of pLM-Guided Data Augmentation for Protein Sequence Prediction Tasks · 2026 · DOI
  • However, engineering carbonic anhydrases to maintain stability under harsh industrial process conditions remains a key challenge, and sequence-to-function datasets compatible with machine learning to inform forward engineering are lacking.

    Machine learning guided cell-free expression maps the biochemical landscape of carbonic anhydrase · 2026 · DOI

Most-cited papers in Machine Learning in Bioinformatics

Most recent work

Find a gap in your own Machine Learning in Bioinformatics sub-topic

This page shows what the Machine Learning in Bioinformatics literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.

Open the Research Gap Finder →

Related topics in Biochemistry, Genetics and Molecular Biology

37 open questions have been extracted from the limitations and future-work passages of 268 Machine Learning in Bioinformatics papers in our library. Each one below links back to the study that raised it, so you can read the original claim in context.

Tools for your next paper

Compare the categoryHonest roundups of the AI research tools, ours listed alongside the alternatives.

Command palette

Jump anywhere, run any action.