Open research questions in Bioinformatics and Genomic Networks
49 unresolved questions extracted from the limitations and future-work sections of 272 Bioinformatics and Genomic Networks papers in our library. Each links back to the study that raised it.
What the literature leaves open
Omics-driven systems pharmacology provides a promising strategy to overcome these limitations, but generative AI tools specifically designed for systems pharmacology-oriented drug design remain scarce.
GEM-GPT Enables Personalized Cell Type-Resolved Therapeutic Design for Systems Pharmacology · 2026 · DOIAlthough AI agents can perform therapeutic analyses, existing systems often fail to preserve biological context over long workflows, verify intermediate computational steps, or reconcile conflicting evidence across datasets and literature.
These findings support a previously underrecognized association between PANX1 and inflammatory tumor biology and suggest its potential value as a biomarker candidate, while further mechanistic and translational validation is required.
Machine learning prioritization identifies PANX1 as an inflammation-associated candidate regulator in lung adenocarcinoma · 2026 · DOIThese results support further study of biologically structured latent-response prediction, while the lower gene-space accuracy and sensitivity to sparse graph neighborhoods limit the scope of the present conclusions.
Stable-Shift: Biologically Structured Prediction of Transcriptional Responses to Unseen Gene Perturbations · 2026Finally, we evaluate the application of CLASPP in predicting PTMs in different model organisms and experimentally validated ubiquitination sites in the understudied DCLK3 kinase.
This pattern extends to 248,191 downstream papers that consume structural knowledge, where engagement with genes lacking experimental structures and with understudied human genes increased since 2021.
AI predictions and the expansion of scientific frontiers: Evidence from structural biology · 2026 · DOIFinally, reliance on bulk transcriptomic data does not account for tumor microen- vironment composition or intra-tumoral heterogeneity, which are increasingly recognized as critical determi- nants of HCC progression and therapeutic response.
Integrated transcriptomic and network biology analysis reveals druggable hub genes as candidate therapeutic targets in hepatocellular carcinoma · 2026 · DOIincreased surveillance. interventions or for therapeutic The effectiveness of DTs depends on their predictive capability (Fuller et al., 2020). Leveraging real-time data and historical health information helps DTs forecast various factors such as disease progression, outcomes of different treatment options, or early signs of health deterioration. This predictive capability is particularly valuable in the management of chronic diseases, influence disease where early progression. DTs can be used not only for optimizing disease treatments, but also to suggest strategies for preventing disease onset and mitigating its impact. Their predictive power comes from modeling disease trajectories, i.e., mapping a disease’s progression over time, including its onset, development and chronicity, when relevant, starting from clinical records, multiomics data and real-time health metrics. Similar data-driven intervention can significantly they offer approaches have also been applied to predict disease risk and evolution using longitudinal health data and ML (Lasko et al., 2025). Therefore, for medical professionals to proactively adjust treatment plans, and anticipate significant turning points in the disease’s course. They can also be used to provide alerts for early indicators of health deterioration for chronic diseases, whose trajectory frequently includes periods of remission or relapse. the potential Moreover, the DTs framework can help in the creation of a learning healthcare system where continuous feedback from DTs informs clinical decisions and enhances healthcare delivery. This iterative process not only improves individual patient care, but also the broader knowledge base, driving key contributes advancements in medical research and clinical practice (Kamel Boulos and Zhang, 2021). to Potential and existing applications of DTs or DT-like models in healthcare There are three main uses that DTs could have in healthcare, namely, disease management and treatment optimization, virtual clinical trials, and patient education. Personalized medicine and disease management are a few of the areas in which DTs are gaining recognition for their transformative potential in the healthcare industry. Several existing systems integrate multi-omics data with clinical and lifestyle information to create a more complete and individualized health profile that can serve as a basis for in silico disease modeling. These models can form the basis for clinical decision support systems, providing predictions of responses to the different treatment options, a highly valuable tool for clinicians who can tailor interventions based on the unique characteristics of each patient. In addition, DTs enable customized interventions by utilizing real-time data to offer precise and dynamic insights into patient health; this involves continuous mapping of the physical/digital counterparts. Virtual representations play an important role in the prevention and management of chronic disease. The constant monitoring of physiological data allows DTs to identify early signs of disease onset but also aggravation, which is crucial for timely intervention. In the context of diabetes management, DTs have demonstrated the capacity to predict glucose fluctuations by integrating data on dietary intake, physical activity or medication adherence, making them a powerful tool to help clinicians with better management of the disease and preventing complications (Shamanna et al., 2021). Recently, this has led to the design of DT systems that administer insulin through real-time monitoring of patient glucose levels (Thamotharan et al., 2023). Other studies in cardiovascular health demonstrate how DTs can monitor vital or relevant cardiovascular signs (resting heart rate, heart rate variability, blood pressure) to detect abnormalities at an early stage and suggest preventive measures (Coorey et al., 2022). Specifically in oncology, advanced computational frameworks have long been used for modeling tumor growth and predicting how a tumor will respond to different therapeutic interventions (Benzekry et al., 2014), leading to optimization of cancer patient treatment plans by improving drug effectiveness and minimizing side effects (Bruynseels, Santoni de Sio and van den Hoven, 2018).
Multilayer network approaches to omics data integration in digital twins for cancer research · 2026 · DOIthe broader that could potentially slow down application of DTs in clinical settings. In the literature, it has been argued that hypothesis-driven generative models, which generate data based on a set of underlying assumptions of the processes that generate the observed data, more particularly multi-scale modeling, are essential to boost the clinical accuracy of DTs. The transformative potential of DTs in healthcare has been explored by emphasizing their capability to simulate complex interdependent biological processes across multiple scales. The integration of generative models with extensive datasets can deliver scenario-based approaches to explore diverse therapeutic strategies. It is an excellent strategy to support dynamic clinical decision-making. This method not only leverages advancements in data science to it also incorporates insights from complex systems, quantitative biology and digital medicine to improve patient care (De Domenico et al., 2025).
Multilayer network approaches to omics data integration in digital twins for cancer research · 2026 · DOIThis study addresses a critical challenge in computational drug repurposing for Alzheimer’s Disease: the need for methods that can systematically discover mechanistically novel and interpretable therapeutic hypotheses. We presented a framework that achieves this by uniquely integrating biologically-informed GNN embeddings with a MAP-Elites quality-diversity search within the TPOT2 AutoML pipeline. Our primary contribution is the use of a biomedical knowledge graph (AlzKB) to guide the evolutionary search, defining novelty based on an embedding’s distance to known AD-related entities. The identification of diverse candidates such as Exemestane and Felodipine from underrepresented regions of the feature space provides supporting evidence for the feasibility of this novelty-guided approach as a hypothesis-generation framework. By fusing graph-based biological priors with a diversity-aware evolutionary search, our work presents a methodological proof-ofconcept for integrating biological knowledge into AutoML-driven drug repurposing. Future work will focus on enriching the biological data foundation by incorporating multi-omics datasets and exploring more advanced GNN architectures. We will also investigate more dynamic AutoML search paradigms using reinforcement learning and the integration of Large Language Models (LLMs) to automatically update the knowledge base. Ultimately, the most critical next step is the experimental validation of our predicted drug candidates through in vitro and in vivo studies, and the application of this framework to other complex neurodegenerative diseases, such as Parkinson’s Disease, Amyotrophic Lateral Sclerosis (ALS), and Huntington’s Disease. Shao et al.
A biology-based quality-diversity algorithm for drug repurposing in Alzheimer’s disease using automated machine learning · 2026 · DOIThe true test of GETgene-AI's utility will lie not in retrospective alignment with known biology but in its prospective ability to predict novel, therapeutically actionable targets that translate into clinical benefit.
The ablation study section is incomplete in the excerpt (ends with 'The full CSGNN m'), making a comprehensive assessment of individual component contributions impossible.
CSGNN's performance on E. coli shows slightly lower AUC and MCC compared to HCNS, indicating potential limitations in handling stronger class imbalance across different species.
Metastasis of cancer remains a significant challenge in oncology and accounts for the majority of cancer-related fatalities globally. Traditional methods for predicting metastasis often fail to incorporate complex multi-omics data and to understand the intricate molecular interactions that promote tumor growth. In this study, we created a new framework that combines multi-omics data like mRNA expression, DNA methylation, gene mutation, and copy number alteration (CNA) with Graph Convolutional Networks (GCNs) and gene embeddings from large language models (LLM) to make pan-cancer metastasis prediction integrates more molecular-level omics characteristics with LLM- derived semantic representations of genes. This lets it topological and contextual biological see both proposed model accurate.The Page | 29 Journal of Hunan University (Natural Sciences) Vol. 53 No. 2, February 2026 relationships, which gives a fuller picture of how tumors act. The proposed graph-based GCN (Omics + LLM) model is a lot better than all of the non-graph baselines, like Transformer, MLP, and Random Forest, when you compare them. It has a 97.37% accuracy, a 97.06% F1- score, and a 99.72% AUC, which is a big improvement over older methods. This shows that adding CNA features to other omics types, as well as PPI- network topology and LLM-based embeddings, makes it easier to understand and predict biological processes. Adding more omics types (like proteomics and metabolomics) and real-time clinical data to this framework will make it better in the future.This will help find problems early and give each patient the best care.Another important goal will be to make it easier to apply what we’ve learned to different datasets and groups of people. Finally, we’ll look into how to use biomedical LLMs with attention and reinforcement learning systems to model how genes interact with each other over time and make predictions about metastasis that are even more useful in the clinic.
SEMO-GCN: Semantic Enhanced Multi-Omics Graph Representation Learning for Pan-Cancer Metastasis Identification · 2026 · DOIPartial correlation networks were not measured at the patient level, but LIONESS networks were constructed [7]. The AUC is not strictly dependent on edge strength, but on which genes are included within modules influenced by morphological inputs. Increasing the number of nodes or genes lowers network resolution. STRING database connections were attempted; however, connections among many Moffitt genes were incomplete or absent, which is expected for sparse or poorly annotated regulatory interactions [8,9]. The globally constructed net- works are not subtype-specific at the patient level but can characterize cohort-level variance related to tumor progression. Nodes were not constructed from morphology and replication timing directly; instead, these modalities influenced edge weights and module scoring.
Personalized Morphology, Replication Timing, and RNA based Gene Expression Networks for Basal-like and Classical subtyping genes in Pancreatic Adenocarcinoma · 2026 · DOIBy constructing a high-resolution knowledge map, we identify underexplored thematic intersections suggested by publication trends, thereby providing a data-driven roadmap to navigate future research priorities.
Mapping the knowledge landscape of planarian regeneration: a century of bibliometric insights · 2026 · DOIAbstract Introduction The complexity of T follicular helper (Tfh) cell populations and inconsistent findings across experimental systems have fueled ongoing debate regarding Tfh differentiation.
A unified network systems approach uncovers a core novel program underlying T follicular helper cell differentiation 2260705 · 2026 · DOIWe found a core Tfh set which is conserved across humans and mice addressing a fundamental open question in the field, as different systems (in-vitro, ex-vivo and in-vivo) have led to discovery of varying components.
A unified network systems approach uncovers a core novel program underlying T follicular helper cell differentiation 2260705 · 2026 · DOIThese features are widely used in image analysis, but their application to biological networks has remained limited because cellular networks are sparse, irregular graphs rather than regular pixel grids.
Textural features for pathway-level representation of omics data in biological networks · 2026 · DOIHowever, the molecular regulators linking inflammatory signaling with tumor biology in lung adenocarcinoma (LUAD) remain incompletely defined.
Machine learning prioritization identifies PANX1 as an inflammation-associated candidate regulator in lung adenocarcinoma · 2026 · DOIThere is a lack of standardized, decision-oriented benchmarks that test whether computational models can generalize therapeutic hypotheses across diseases in ways that reflect real-world pharmaceutical investment decision making.
Clinical Trial and Ontology-Derived Positive and Negative Benchmark Datasets for Drug Repurposing Across Rare Diseases · 2026 · DOIThese scenarios trigger distribution shifts arising from heterogeneous sequencing platforms, distinct tissue microenvironments, and metastatic evolution--problems rarely addressed by existing methods.
CSGDA: A Cell State-Guided Graph Domain Adaptation Network for Single-Cell Drug Response Prediction · 2026 · DOIEven though sampling variability frequently compromises findings, particularly in small cohorts, the reproducibility and generalizability of the derived latent structures are insufficiently evaluated.
Meanwhile, in many biological fields, core regulatory genes have been extensively studied, leading to the establishment of small-scale gene regulatory networks, and novel genes connected to these networks remain to be identified.
Expanding gene regulatory networks from transcriptome data through graphical modeling with heterogeneous priors · 2026 · DOIAdditional analyses show that semantic short-path encoding contributes most to performance, while mechanism-context augmentation improves robustness under sparse evidence and strengthens Gene Ontology functional agreement.
CAREPath: Semantic Context-Aware Reasoning Paths with Mechanism-Augmented Embeddings for Drug Repurposing · 2026 · DOI
Most-cited papers in Bioinformatics and Genomic Networks
- Differential network biology · Molecular Systems Biology · 2012 · 683 citations
- The STRING database in 2025: protein networks with directionality of regulation · Nucleic Acids Research · 2024 · 582 citations
- Nonlinear dynamics of multi-omics profiles during human aging · Nature Aging · 2024 · 406 citations
- WebGestalt 2024: faster gene set analysis and new support for metabolomics and multi-omics · Nucleic Acids Research · 2024 · 380 citations
- OmicShare tools: A zero‐code interactive online platform for biological data analysis and visualization · iMeta · 2024 · 262 citations
- Emerging trends and hot topics in the application of multi-omics in drug discovery: A bibliometric and visualized study · Current Pharmaceutical Analysis · 2024 · 147 citations
- Visualizing set relationships: EVenn's comprehensive approach to Venn diagrams · iMeta · 2024 · 136 citations
- Harmonizome 3.0: integrated knowledge about genes and proteins from diverse multi-omics resources · Nucleic Acids Research · 2024 · 128 citations
- DeepKEGG: a multi-omics data integration framework with biological insights for cancer recurrence prediction and biomarker discovery · Briefings in Bioinformatics · 2024 · 109 citations
- Protein function prediction as approximate semantic entailment · Nature Machine Intelligence · 2024 · 108 citations
Most recent work
- FlyPredictome: A structural atlas of predicted protein-protein interactions in Drosophila · bioRxiv · 2026
- DeepISO: deep learning-powered prediction of protein–protein interaction rewiring generated by alternative splicing · Genome Biology · 2026
- SMODA: Interpretable Multimodal Omics Integration for Disease Classification and Subtype Discovery via Heterogeneous Transfer Learning · Analytical Chemistry · 2026
- A Context-Specific, Literature-Supported Framework for Validating Stress Response Differentially Expressed Gene Sets · bioRxiv · 2026
- Interpreting Omics Data Analysis with Large Language Models for Disease Target and Drug Discovery · bioRxiv · 2026
- Generating Joint Transcriptomic and Morphological Responses to Drug Perturbations via Rectified Flow · bioRxiv · 2026
- Personalized Morphology, Replication Timing, and RNA based Gene Expression Networks for Basal-like and Classical subtyping genes in Pancreatic Adenocarcinoma · 2026
- A CSGNN model-based method for essential protein identification · Frontiers in Bioinformatics · 2026
- Gene Network Enrichment Analysis and Its Application to Explore Enriched Immune Disease Pathways for Gene Network of Acute Myeloid Leukemia Cell Lines · Journal of Computational Biology · 2026
- SEMO-GCN: Semantic Enhanced Multi-Omics Graph Representation Learning for Pan-Cancer Metastasis Identification · Journal of Hunan University Natural Sciences · 2026
Find a gap in your own Bioinformatics and Genomic Networks sub-topic
This page shows what the Bioinformatics and Genomic Networks literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →