Open research questions in Computational Drug Discovery Methods
126 unresolved questions extracted from the limitations and future-work sections of 525 Computational Drug Discovery Methods papers in our library. Each links back to the study that raised it.
What the literature leaves open
deep learning for prediction, lacks prediction with drug compatibility molecular representation integration with drug and synergy evaluation and prediction synergy analysis Vamathevan et al., 2019 AI…
PredictRx: AI based decision support tool for molecular screening for breast cancer drug recommendation · 2026 · DOIRitu Chauhan1,2, Neha Pandey1 and Megat F. Zuhairi2* 1Artificial Intelligence and IoT Lab, Centre for Computational Biology and Bioinformatics, Amity University, Noida, India, 2Malaysian Institute of Information Technology, Universiti Kuala Lumpur, Kuala Lumpur, Malaysia Introduction: Breast cancer remains one of the leading causes of cancerrelated mortality rate worldwide, and the identification of effective drug combinations is an essential requirement in pharmaceutical research. The integration of Artificial Intelligence (AI) in processing large volumes of chemical and biological data combines molecular representation, predictive modeling and structured support within a single accessible tool, which accelerates earlystage candidate identification for breast cancer research while promoting reproducibility, transparency and user centered design. Aim: The current research focuses on developing and designing “PredictRx” which is an artificial intelligence based driven decision support tool which tends to benefit healthcare practioners to analyze the combination of drug which can be utilized for breast cancer patients. Methodology: PredictRx was developed using molecular descriptors, physicochemical properties, and drug interaction datasets collected from integrates in total six publicly available biomedical databases. The tool supervised and unsupervised learning techniques to examine the structural similarities between compounds and predict the potential drug interactions for breast cancer. Various machine learning techniques, including Random Forest, Support Vector Machine, Logistic Regression, K-Means Clustering, DBSCAN, and Agglomerative Clustering, to analyse structural similarities and predict potential drug interactions and synergy patterns. Model performance was evaluated using Classification matrix, Silhouette Score, Calinski-Harabasz Index, and Davies-Bouldin Index. The tool was deployed as a browser-accessible web application for real-time interaction and visualization. Result: The results suggests that Random Forest has the highest predictive performance accuracy of 1, and Agglomerative clustering delivered strongest scores (Silhouette Score: 0.6946; Davies-Bouldin Index: 0.2457). The current tool was deployed as a browser accessible web tool with possibility of real time interaction and result visualization. PredictRx is a distinctive easy to use, and interpretable screening tool focused on drug compatibility and synergy analysis. EDA further identified molecular weight, lipophilicity, and structural similarity as important contributors to drug compatibility prediction. Conclusion: PredictRx shows how AI-driven predictive modeling which can speed up molecular screening and early-stage breast cancer medication Frontiers in Artificial Intelligence 01 frontiersin.org Chauhan et al. 10.3389/frai.2026.1881187 discovery. The technology facilitates the effective identification of appropriate drug combinations and offers a scalable foundation for upcoming AI-assisted pharmaceutical research by combining clustering, classification, molecular representation, and visualization into a single interpretable platform.
PredictRx: AI based decision support tool for molecular screening for breast cancer drug recommendation · 2026 · DOIThe workflows were not prospectively evaluated on compounds lacking pre-existing IZ measurements, and whether retrospective enrichment improves experimental hit discovery or reduces screening workload remains to be established.
Continuous Inhibition-Zone Modeling and Binary Classification for Pseudomonas aeruginosa Hit Prioritization: A Retrospective QSAR Evaluation · 2026 · DOIWe demonstrate this capability by using CovSite as a blind, ligand-specific approach that enables iterative, machine-learning-driven covalent inhibitor generation that is impractical with existing tools, establishing a foundation for computationally guided covalent drug discovery for novel and understudied targets.
CovSite: A High-Throughput Blind Covalent Screening Framework for Reactive Site Detection · 2026 · DOIWhile tools exist for de novo biosynthetic gene cluster identification and large-scale unsupervised clustering, dedicated methods for the targeted, hypothesis-driven expansion of user-defined BGC families are lacking.
DiscERN: an automated genome mining tool for the discovery of evolutionarily related natural products · 2026 · DOIWe applied this framework to histone deacetylase 11 (HDAC11), the sole class IV member of the histone deacetylase family and epigenetic regulator implicated in tumor progression and therapy resistance, yet remains chemically underexplored with relatively few inhibitors available.
Abstract A024: FastBindRank, a novel, scalable method for high-fidelity virtual screening of ultra-large chemical libraries idendtifies novel HDAC11 inhibitors · 2026 · DOISeveral near-term methodological developments are likely to strengthen the integrated GNN/core-hopping framework proposed here. Active-learning loops that iteratively retrain the GNN on newly generated experimental data from each design cycle should progressively sharpen mechanistic-class discrimination with far fewer total compounds synthesized than a static-model campaign would require. Incorporation of MD-refined, rather than purely docked, poses into GNN training data feeding the model three-dimensional interaction fingerprints time-averaged over simulation trajectories rather than single static docking poses — should improve the model's ability to distinguish the subtle H12-engagement differences that separate full from partial agonists. Explainable-AI techniques applied to trained GNNs, which can highlight which substructural motifs or atom-level contributions drive a given mechanistic-class prediction, would directly address the interpretability limitation noted above and could, in principle, surface previously unappreciated pharmacophoric features relevant to biased agonism. Finally, tighter integration of generative de novo design with the core-hopping paradigm described in Section 6 using reinforcement-learning or diffusionbased generative models constrained simultaneously by predicted affinity, mechanistic-class probability, synthetic accessibility, and PBPK-predicted pharmacokinetics represents a natural extension of the multiparameter-optimization generative frameworks already demonstrated for other CNS- and oncologytargeted programs [38,40,41] to the specific, mechanistically nuanced requirements of PPAR-γ modulator discovery.
GRAPH NEURAL NETWORK-GUIDED VIRTUAL SCREENING AND CORE-HOPPING-BASED DESIGN OF NOVEL PPAR-γ MODULATORS · 2026 · DOIEventually, the in silico screen does not account for pharmacokinetic properties such as CNS penetration, metabolic stability, or plasma protein binding; integration of ADMET prediction models would improve the translational relevance of the candidate list.
Machine-learning-based pIC50 prediction identifies novel acetylcholinesterase inhibitors among FDA-approved drugs · 2026 · DOINotably, processed foods are often packaged in PVC materials containing plasti- cizers, which may serve as an underrecognized environmental risk factor for GC progression.
Screening of core targets for Di(2-ethylhexyl) Phthalate-related gastric cancer based on machine learning, molecular docking, and SHAP analysis · 2026 · DOIWhile prior studies have focused on learning molecular representations for individual components, modeling how multiple components and their ratios jointly influence LNP performance remains underexplored.
One limitation of our study is that our tested acquisition functions were static and did not adapt to changes in availability of compounds in the untested pool. Our study is also limited by the focus of AL on a single molecular property.
Future work could explore wider hyperparameter grids for t-SNE and UMAP, the application of additional molecular representations beyond ECFP4 finger- prints, and the extension of this pipeline to even larger datasets approaching the scale of the complete drug-like chemical space.
No AI framework matched the scope or depth of the original studies, results varied across multiple runs of the same framework with the same prompt, and we documented cases of severe hallucinations in final reports, gaps in literature coverage, and overconfident conclusions.
Real Science Is Harder Than Benchmarks: Evaluating Advanced AI Frameworks on Published Studies. I. Uncertainty Quantification, ML on Therapeutic Data Commons, and Agent-Based Modeling · 2026 · DOIOverall, these results indicate that the iterative refinement procedure was generally robust, with limited evidence of degeneration except for increased historical reuse in the llama-based variant. In addition, the increase in predicted pKi was accompanied by declining validity and QED together with increasing logP, indicating that affinity-oriented optimiza- tion alone is insufficient as a sole design objective.
Target-aware molecule SMILES generation using a large language model with retrieval-augmented generation, multi-turn memory, and a predictive model · 2026 · DOIStill, methods for predicting further PK‐shape‐defining parameters like absorption rate constant (k a ), central and peripheral distribution volume (V c , V p ), and intercompartmental clearance (Q) remain underexplored.
Breaking Barriers Between Species: Integrated Multi‐Species <scp>PK</scp> Models Improve Human <scp>PK</scp> Predictions · 2026 · DOIYet, empirical evidence systematically characterizing the structural and evolutionary dynamics of this phenomenon in AI-driven drug discovery remains scarce.
Structural characteristics and evolutionary trajectories of knowledge recombination in the field of AI-driven drug discovery · 2026 · DOINotably, CRUSH-derived fragments access exclusive regions of chemical space that remain unexplored by existing approaches, a result that is consistent across both low- and high-resolution molecular fingerprint representations.
CRUSH—Cleavage Rules Using SMIRKS Heuristics: an enhanced molecular fragmentation algorithm · 2026 · DOIAlthough docking and related approaches that explicitly account for protein-ligand interactions have been developed and refined over several decades, achieving both reliable protein-aware interaction modeling and computational scalability remains an open challenge, particularly for ultra-large chemical spaces.
This study employs some basic physicochemical properties to predictive and rank the kidney targeted drug candidates using computational models, Random Forest and RWNN, together with TOPSIS for multi criteria decision making. It showcases the models predictive potential of the param- eters such as polarizability, and binding along with the framework for the ranking of drug candidates. The study is however confined to some selected drugs and topologi- cal descriptors, and the resultant structural and pharmaco- kinetic variations which may go unaccounted in larger and more diverse chemical libraries. Though the models have shown reliability and generalization regarding the dataset used, experimental prediction validation is an outstanding cornerstone regarding the predictive worth of the models.
Computational intelligence-based QSPR modeling and decision-making for anti-kidney cancer drugs evaluation using machine learning algorithms · 2026 · DOIFurther studies could enhance this computational approach through the inclusion of larger and more diverse datasets, including additional anti-cancer drugs and multi-target pre- dictions. We may make predictions even more precise and learn more about how molecules interact in more complex ways by combining advanced deep learning structures, including graph neural networks. Moreover, incorporat- ing experimental validation with in simulation predictions would enhance the certainty of the suggested strategy. This method can also be utilised to examine the interactions of various medications, so enhancing the efficacy and effi- ciency of combination therapy for kidney cancer and other cancers, ultimately enhancing drug development workflows. Author contributions All the authors Wakeel Ahmed, Maryam Alvi and Shahid Zaman have equally contributed to this manuscript in all stages, from conceptualization to the write-up of final draft. Funding There is no funding to support this article. Data availability All data generated or analysed during this study are included in this published article.
Computational intelligence-based QSPR modeling and decision-making for anti-kidney cancer drugs evaluation using machine learning algorithms · 2026 · DOIOverall, CMA-DTI provides a practical framework for multimodal DTI prediction, and future work will focus on incorporating more strictly curated binding data and explicit three-dimensional structural information to further improve generalization and interpretability.
CMA-DTI: a cross-modal fusion and attentive interaction network for interpretable drug-target interaction prediction · 2026 · DOIFuture work will focus on more strictly curated binding data, broader transfer and cold-start evaluation across independently curated datasets, controlled label-noise sensitivity analysis, and the integration of explicit three-dimensional structural information.
CMA-DTI: a cross-modal fusion and attentive interaction network for interpretable drug-target interaction prediction · 2026 · DOIOur protein list, derived from predictions, highlights glioma-relevant targets but is inherently incomplete, similar to databases like UniProt or GeneCards, as each captures only a partial view of glioma biology. While this list serves as one of the curated gold standards for our analysis, incorporating known treatment targets in future studies could provide a more comprehensive benchmark. Limitations of this study include the arbitrary rank cutoffs, which may exclude moderately ranked targets that overlap meaningfully with gold standard libraries, and the use of the Jaccard coefficient, a binary metric that overlooks relative ranks or prediction scores. Additionally, our focus on glioma leaves the robustness of this approach across other indications underexplored, particularly for diseases with fewer validated targets. Finally, the analyses may bias toward frequently predicted top targets, underrepresenting less common targets with potential therapeutic value. To address these limitations, future studies will integrate score-based cutoffs, and consider a broader range of rank and score distributions. Xu et al.
Multiscale analysis and optimal glioma therapeutic candidate discovery using the CANDO platform · 2026 · DOIimproving predictive reliability, interpretability, and translational impact. is essential these A structured summary of major challenges and proposed solutions is provided in (Table 7).
Quantitative Structure–Activity Relationship Modeling of Benzothiazole Derivatives: Methodological Advances, Biological Applications, and Future Perspectives · 2026 · DOIIntegration of publicly available bioactivity data from curated databases such as ChEMBL and PubChem can expand dataset size and chemical diversity [63,64]. However, careful data curation remains critical to prevent noise introduction [21,22].
Quantitative Structure–Activity Relationship Modeling of Benzothiazole Derivatives: Methodological Advances, Biological Applications, and Future Perspectives · 2026 · DOI
Most-cited papers in Computational Drug Discovery Methods
- ProTox 3.0: a webserver for the prediction of toxicity of chemicals · Nucleic Acids Research · 2024 · 1,260 citations
- PubChem 2025 update · Nucleic Acids Research · 2024 · 1,066 citations
- ADMETlab 3.0: an updated comprehensive online ADMET prediction platform enhanced with broader coverage, improved performance, API functionality and decision support · Nucleic Acids Research · 2024 · 975 citations
- SwissDock 2024: major enhancements for small-molecule docking with Attracting Cavities and AutoDock Vina · Nucleic Acids Research · 2024 · 482 citations
- Deep-PK: deep learning for small molecule pharmacokinetic and toxicity prediction · Nucleic Acids Research · 2024 · 264 citations
- Reinvent 4: Modern AI–driven generative molecule design · Journal of Cheminformatics · 2024 · 248 citations
- ADMET-AI: a machine learning ADMET platform for evaluation of large-scale chemical libraries · Bioinformatics · 2024 · 247 citations
- An artificial intelligence accelerated virtual screening platform for drug discovery · Nature Communications · 2024 · 200 citations
- Open Targets Platform: facilitating therapeutic hypotheses building in drug discovery · Nucleic Acids Research · 2024 · 183 citations
- Structure-based drug design with equivariant diffusion models · Nature Computational Science · 2024 · 180 citations
Most recent work
- Leveraging machine learning models in evaluating ADMET properties for drug discovery and development · ADMET and DMPK · 2026
- Sparse Autoencoders Reveal Interpretable Features in Single-Cell Foundation Models · bioRxiv · 2026
- Artificial intelligence in drug discovery from advanced molecular representation to pipeline applications · Frontiers in Bioinformatics · 2026
- A computational pipeline combining machine learning and molecular simulations identifies repurposed PI3Kα inhibitors from FDA-approved drugs · Artificial Intelligence Chemistry · 2026
- Programming Biomolecular Interactions with All-Atom Generative Model · bioRxiv · 2026
- TCMNet: an AI-driven strategy for optimizing traditional Chinese medicine · Chinese Medicine · 2026
- Recent advances in AI-driven pKa prediction for proteins and small molecules · Current Opinion in Structural Biology · 2026
- Machine Learning Approach to Anticancer Activity Prediction of Transition-Metal Complexes Based on a Large-Scale Experimental Database · Journal of Medicinal Chemistry · 2026
- A Primer Exchange Reaction-Based Multidimensional AND-Gated Molecular Classifier · Analytical Chemistry · 2026
- Artificial intelligence-based screening of phytochemicals for targeted cancer therapy · Natural Products and Bioprospecting · 2026
Find a gap in your own Computational Drug Discovery Methods sub-topic
This page shows what the Computational Drug Discovery Methods literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →