Open research questions in Genetics, Bioinformatics, and Biomedical Research
61 unresolved questions extracted from the limitations and future-work sections of 1,028 Genetics, Bioinformatics, and Biomedical Research papers in our library. Each links back to the study that raised it.
What the literature leaves open
The future of crop protection is expected to be shaped by the convergence of advanced biotechnology, artificial intelligence, precision agriculture, and microbiome science. Among the most transformative is CRISPR-Cas gene-editing technology, innovations which offers unprecedented precision in modifying genes associated with pest resistance, disease tolerance, and beneficial microbial interactions (Wang introduction, CRISPR has et al., 2022). Since significantly accelerated plant improvement efforts by enabling targeted genetic modifications without introducing extensive foreign DNA. Researchers are already exploring gene-edited microbial biopesticides capable antimicrobial compounds, improved environmental persistence, and greater effectiveness against resistant pathogens (Halder et al., 2022). producing enhanced its of Table 4: Current Limitations and Proposed Technological Solutions.
The Biopesticide Revolution: AI, Genomics, and Microbiome Engineering on a Global Scale · 2026 · DOIWhile largely driven by ML architecture refinements, these advancements in ML-based protein engineering campaigns have left the impact of data curation underexplored.
EvoSeq-ML: Advancing Data-Centric Machine Learning with Evolutionary-Informed Protein Sequence Representation and Generation · 2026 · DOIThe paper demonstrates that Cysteine clusters with polar amino acids rather than Methionine despite both containing sulfur atoms, suggesting overall physicochemical properties dominate specific chemical features, but does not systematically test this principle across amino acid pairs with shared chemical elements to establish the relative weighting of feature classes.
Unraveling protein secrets: machine learning unveils novel biologically significant associations among amino acids · 2026 · DOIThe multidimensional amino acid descriptors used in hierarchical clustering are not explicitly detailed in this excerpt, making it unclear whether the analysis incorporates structural descriptors (backbone dihedral angles, solvent accessibility) versus purely physicochemical descriptors, limiting reproducibility and comparison with alternative descriptor sets.
Unraveling protein secrets: machine learning unveils novel biologically significant associations among amino acids · 2026 · DOIThe study identifies low-consensus substitutions involving Gly and Pro as high-risk mutations based on consensus thresholds (<0.30), but does not validate these predictions against experimental mutagenesis datasets or provide specificity regarding which protein structure contexts make these substitutions more or less disruptive.
Unraveling protein secrets: machine learning unveils novel biologically significant associations among amino acids · 2026 · DOIThe hierarchical clustering reveals context-dependent roles for Glycine and Proline with variable clustering behavior across subsamples, but the paper does not specify which structural contexts (secondary structure type, protein fold class, or functional region) drive their variable clustering patterns.
Unraveling protein secrets: machine learning unveils novel biologically significant associations among amino acids · 2026 · DOIThe study proposes that cation-π interactions between positively charged amino acids (Arg, His, Lys) and aromatic rings explain their clustering together, but does not quantitatively assess the frequency and strength of these interactions across a systematically sampled protein structure dataset to validate this mechanistic hypothesis.
Unraveling protein secrets: machine learning unveils novel biologically significant associations among amino acids · 2026 · DOIThe consensus clustering approach identifies high-consensus clusters (>0.85) as functionally interchangeable for protein engineering applications (e.g., Trp-Phe-Tyr aromatic triad), but empirical validation through directed mutagenesis experiments within these specific clusters in real proteins is not demonstrated.
Unraveling protein secrets: machine learning unveils novel biologically significant associations among amino acids · 2026 · DOIThe study identifies Methionine's intermediate positioning between branched aliphatic amino acids and other groups in hierarchical clustering, suggesting a bridging role at protein domain boundaries, but lacks experimental validation of this proposed function through domain boundary analysis or structural studies of Methionine occurrence patterns at domain interfaces.
Unraveling protein secrets: machine learning unveils novel biologically significant associations among amino acids · 2026 · DOIBull Math Biol 86(11):135 Pastva S, Park KH, Huvar O, Rozum JC, Albert R (2025) An open problem: Why are motif-avoidant attractors so rare in asynchronous Boolean networks? J Math Biol 91(1):11 Spector R, Harrington HA, Gaffney EA (2026) Persistent homology classifies parameter dependence of patterns in Turing systems.
The distinction between genetic information patterns that produce robust multicellular organisms versus disordered tissue in cancer is proposed but lacks specification of which eigengenes, which intracellular dynamics controlling protein folding, or which evolutionary selection mechanisms fail or become dysregulated in cancer cell proliferation compared to normal development.
The paper suggests that non-random cellular interactions during proliferation produce information gain analogous to amino acid interactions in protein folding, but does not specify which cell-cell signaling pathways, which types of non-random interactions, or quantitative metrics should be used to measure information gain during actual tissue morphogenesis and distinguish it from random proliferation.
The model claims that the 30 billion bits of human genome information must be significantly expanded to form 36 trillion cells, but lacks a detailed calculation or experimental design comparing actual Shannon entropy measurements in genomic sequences versus the information content required for cell-type specification and tissue organization across different developmental stages.
The paper invokes eigengenes and co-expression patterns of genes in adjacent cells as the proposed biological mechanism for dimensional information expansion during brain development, but does not specify which eigengene patterns, which brain tissues, or at which developmental timepoints this mechanism should be investigated to empirically validate the theoretical model.
The paper proposes that information gain through dimensional expansion (1D to 3D) applies to multicellular organism development analogous to protein folding, but provides no quantitative framework for measuring Shannon information content during embryonic development stages or validating the information expansion claim empirically in developmental systems beyond the single Choanoeca flexa example.
In conclusion, through the machine-learning-driven informatics methods, this scientometric analysis offers an objective and comprehensive overview of global AlphaFold research, identifying critical research clusters and hotspots while prospectively pointing out underexplored critical areas.
Artificial intelligence alphafold model for molecular biology and drug discovery: a machine-learning-driven informatics investigation · 2024 · DOIThe text emphasizes that molecular recognition and drug binding are dynamic processes involving conformational changes and 'jiggling' of both target and ligand, and that MD simulations reveal protein dynamics beyond X-ray crystallography. However, it does not specify quantitative thresholds for when protein flexibility significantly impacts docking predictions, benchmark datasets comparing static vs. dynamic docking performance, or guidelines for selecting force fields and simulation timescales for different protein classes in computational drug screening.
The paper advocates targeting protein-protein interaction edges rather than nodes (hubs) to avoid indiscriminate pathway blockade, and discusses both orthosteric and allosteric disruption mechanisms. However, it provides no comparative data on the prevalence of allosteric vs. orthosteric PPI inhibition sites in real networks, or systematic computational screening strategies to predict which disruption mechanism is feasible for a given PPI pair.
Fragment-based drug discovery screening uses high-throughput docking followed by MD simulations with implicit-solvent force fields, then explicit-solvent simulations for ligand optimization. The paper does not specify validation benchmarks for the transition between implicit and explicit solvent protocols, optimal fragment library sizes relative to chemical space coverage, or quantitative criteria for hit-to-lead optimization in diverse therapeutic targets.
The paper discusses cryptic binding sites that form only in ligand-bound structures and are not detectable in ligand-free crystal structures, noting they are mostly located away from the main functional site. However, it lacks specification of the frequency of cryptic sites across different protein classes, criteria for predicting their druggability a priori, or systematic protocols to identify and screen these sites computationally in structure-based drug discovery workflows.
The paper identifies hot-spot residues at protein-protein interaction interfaces as critical for PPI inhibitor design, but notes that experimental determination of hot-spots is tedious and time-consuming. Computational methods for hot-spot prediction are mentioned as developed, yet the paper does not specify validation metrics, accuracy rates, or comparative performance of these computational approaches against experimentally determined hot-spots in diverse PPI networks.
While FISH technology is presented as solving karyotype identification problems, the text does not discuss the scalability or practical application of multi-probe fluorescence labeling across diverse plant species and complex genomic systems.
Molecular Biology Technology in Environmental Biology under the Background of Artificial Intelligence · 2020 · DOIThe Roadmap identifies gaps, opportunities and key actions to help guide research towards removing synthesis as a constraint to the access of any given molecule and achieving 100%…
The computer output does in fact give some ground for conjecture that groups ( a ) a n d (b) are related to manuscript D in the form: D (b) (a) Beyond this, the results are insufficiently consistent to support even a plausible conjecture, and obvi- ously a larger sample needs to be used in any serious attempt to elucidate the relationships of all the extant texts.
The above discussion is necessarily sketchy, but it clearly brings out three important points: (a) an enormous amount of work remains to be done without ever going to multicellular organisms; (b) important problems exist at all levels of complexity; and (c) there is likely to be an increasing demand for large amounts of pure cell components present in the cell in rather small amounts.
Most-cited papers in Genetics, Bioinformatics, and Biomedical Research
- Database resources of the National Center for Biotechnology Information in 2025 · Nucleic Acids Research · 2024 · 472 citations
- Artificial intelligence alphafold model for molecular biology and drug discovery: a machine-learning-driven informatics investigation · Molecular Cancer · 2024 · 177 citations
- Using Systems and Systems Thinking to Unify Biology Education · CBE—Life Sciences Education · 2022 · 55 citations
- AI protein-prediction tool AlphaFold3 is now more open · Nature · 2024 · 49 citations
- Insight from Biology Program Learning Outcomes: Implications for Teaching, Learning, and Assessment · CBE—Life Sciences Education · 2023 · 23 citations
- Mapping research on biochemistry education: A bibliometric analysis · Biochemistry and Molecular Biology Education · 2022 · 20 citations
- A framework for understanding the characteristics of complexity in biology · International Journal of STEM Education · 2016 · 20 citations
- Project K: "The Complete Solution of E. Coli " · Perspectives in biology and medicine · 1973 · 17 citations
- Punnett Squares or Protein Production? The Expert–Novice Divide for Conceptions of Genes and Gene Expression · CBE—Life Sciences Education · 2021 · 14 citations
- A feeling for the (micro)organism? Yeastiness, organism agnosticism and whole genome synthesis · New Genetics and Society · 2020 · 13 citations
Most recent work
- Unraveling protein secrets: machine learning unveils novel biologically significant associations among amino acids · Network Modeling Analysis in Health Informatics and Bioinformatics · 2026
- Global Change Biology Communications: Expanding the Conversation on Global Change Biology · Global Change Biology Communications · 2026
- Celebrating 25 Years: Promoting and Disseminating Computational Biophysics and Chemistry · Journal of Computational Biophysics and Chemistry · 2026
- When Genes Meet Screens: Testing the Effect of Bioinformatics-based Activities on the Achievement of Eleventh Grade Palestinian Biology Students · Journal of Science Education and Technology · 2026
- How Does Genetic Information Enable Life? · Bulletin of Mathematical Biology · 2026
- Problems, Progress and Perspectives in Mathematical and Computational Biology · Bulletin of Mathematical Biology · 2026
- Mathematics in Bioinformatics: Foundations, Methods, and Emerging Directions · International Journal of Computer Applications Technology and Research · 2026
- Program Logic for Genomics Education · Zenodo (CERN European Organization for Nuclear Research) · 2026
- Student-Driven Microbiome Exploration: A Low-Cost 16S rRNA Sequencing Curriculum for Undergraduate Biology Education · bioRxiv · 2026
- Genetic Engineering with Quantum Circuits: creating codes and studying BioBloQu genetic elements · bioRxiv · 2026
Find a gap in your own Genetics, Bioinformatics, and Biomedical Research sub-topic
This page shows what the Genetics, Bioinformatics, and Biomedical Research literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →