Open research questions in Radiomics and Machine Learning in Medical Imaging
67 unresolved questions extracted from the limitations and future-work sections of 332 Radiomics and Machine Learning in Medical Imaging papers in our library. Each links back to the study that raised it.
What the literature leaves open
These findings indicate that longitudinal breast MRI-derived biomarkers warrant further investigation for early treatment-response assessment; however, the present models remain exploratory and require independent external validation before any clinical application.
Longitudinal Breast MRI for Early Treatment-Response Modeling: A Comparative Study of Handcrafted Radiomics and Frozen Deep Image Embeddings for pCR Prediction · 2026 · DOIHowever, applicability and the suboptimal ratio of NPC: non‐NPC patients remain the weaknesses of existing studies, warranting further investigation with cohorts that better reflect the screening scenario.
Systematic Review of <scp>AI</scp> ‐Driven Applications for Screening Nasopharyngeal Carcinoma Using <scp>MRI</scp> · 2026 · DOIABSTRACT Background Nasopharyngeal carcinoma (NPC) can be detected early on MRI, but adoption for screening is limited by a shortage of experienced specialists.
Systematic Review of <scp>AI</scp> ‐Driven Applications for Screening Nasopharyngeal Carcinoma Using <scp>MRI</scp> · 2026 · DOIMost prior urinary-metabolite work, including our 2024 creatine-riboside and N-acetylneuraminic-acid report, used case-control designs without a locked, integrated model, and--even where two cohorts were analyzed--lacked prespecified independent validation or TRIPOD+AI-compliant reporting; an interpretable urinary index integrating an expanded metabolite panel with clinical variables under these standards had not been described.
The urinary-metabolite-based lung cancer index (uLCI): an interpretable machine-learning risk model for early-stage disease · 2026 · DOIFuture work should focus on prospective validation of thresholds that prioritize clinical utility, such as minimizing false positives for high- stakes decisions. Last, despite the promising AUC of our fusion model, we acknowledged a key limitation regarding its clinical applicability, particu- larly in the external test cohort. However, the prior study was limited by a smaller sample size (the sample sizes were 151, 159, and 193, respectively), and a lack of habitat analysis.
MRI-based quantification of intratumoral heterogeneity for predicting recurrence risk in ER+/HER2− breast cancer · 2026 · DOIAutomated CT-derived body composition analysis has emerged as an objective marker of patient physiological reserve, but its value in prognostication in TACE patients is insufficiently studied.
Fully automated CT-based quantitative body composition analysis for predicting survival in patients with HCC undergoing TACE: a dual-cohort study · 2026 · DOIWe sought to evaluate whether a new breast MRI image-based artificial intelligence (AI) tool, with or without standard clinicopathologic (CP) factors, can (1) predict high 21-gene score (≥26) for guiding adjuvant chemotherapy and (2) prognosticate 5-year recurrence risk better than the 21-gene assay.
Predictive and prognostic value of MRI radiomics in patients with hormone-receptor–positive, HER2-negative breast cancer: A multicenter study. · 2026 · DOISeveral limitations should be acknowledged. First, the data- set originated from a single curated academic educational archive and likely overrepresented classical, teaching-ori- ented entities while underrepresenting atypical, borderline, or consultation-level cases, which may limit generaliz- ability to routine OMFP practice. In addition, the dataset contained unequal representation of diagnostic entities, with certain lesions occurring multiple times while others were represented by single cases; this imbalance may have influenced category-level accuracy estimates. Although the sample spanned 86 unique diagnoses, it remains modest. Second, the study used an intentionally text-only, zero-shot design. In routine practice, oral and maxillofacial patholo- gists integrate histopathology with gross findings, imaging, and clinical history, and recent work suggests that adding images or contextual data can improve AI performance [35, 42]. Thus, the present estimates reflect a constrained benchmark of text-only interpretation rather than full clini- copathologic diagnosis. Third, some included archival cases represented entities in which histopathologic findings were present but were not independently determinative of the final diagnosis, such as lesions requiring systemic, endo- crine, metabolic, or clinicoradiologic correlation. Although such cases were intentionally classified as CD, their inclu- sion may have broadened that category and should be considered when interpreting category-level and HS/CD- stratified findings. Fourth, although HS/CD classification was performed independently by two oral and maxillofa- cial pathologists with consensus resolution, this operational framework was developed for the present study and was not externally validated; therefore, some subjectivity in classifi- cation remains inevitable. Fifth, repeat scoring with LLMs was performed by a single observer without independent multi-rater validation or formal inter-rater reliability test- ing. Sixth, the models were accessed through publicly avail- able interfaces in their default settings, so full control over backend parameters, hidden system settings, model updates, and session-level variability was not possible. Training data cutoffs were also unknown, and some overlap with publicly available educational material cannot be excluded. Finally, we evaluated only three general-purpose models and did not include domain-adapted or multimodal systems. Alternative architectures, prompting strategies, or human-in-the-loop workflows may yield different performance profiles. Despite these limitations, the study provides a controlled benchmark of contemporary LLM performance on free- text histopathologic narratives in OMFP. The relatively 1 3Head and Neck Pathology (2026) 20:65 65 Page 12 of 14 high performance observed in several HS categories sug- gests potential utility in supervised educational, informatics, and retrospective data-structuring workflows. Future work should prioritize external validation, multimodal integra- tion, domain adaptation, and human-in-the-loop evaluation [38, 43–45].
Diagnostic Performance of Contemporary Large Language Models on Free-Text Histopathologic Descriptions in Oral and Maxillofacial Pathology · 2026 · DOIof The inherent “black-box” opacity of advanced ML models often restricts their clinical translation. Utilizing the SHAP framework helped relate model predictions to radiomic features with plausible biological relevance. In our model, IMAT and SM features accounted for nearly 90% of the predictive contribution. The topranking SM_wav_HH_glszm_ SizeZoneNonUniformityNormalized, characterizes spatial textural heterogeneity.
Machine learning based on body composition radiomics for predicting early recurrence in colorectal cancer: a multicenter study · 2026 · DOIThe radiology AI is not merely new technology; it's a disruption to Precision Radiology. As these and other examples in this review illustrated, the AI impact stretches across the entire imaging value chain, from raw data acquisition/derivation through to complex interpretation of radiomic features allowing for "virtual biopsies" [72]. Whilst concern for AI's inevitability overtaking all radiologists was very much a topic of early discussion, the collective mindset has cert ainly shifted towards Augmented Intelligence. In this scenario, the radiologist continues to be the center of clinical integration, utilizing IA tools to dehumanize repetitive low-energy tasks and focusing their attention on complex diagnostic issue-solving problems and patient engagement [73]. The future achievement of AI in this domain will be based on our ability to address the "black box" nature of algorithms via Explainable AI (XAI), and to establish that models are trained on a diverse, multi-institutional dataset to eradicate demographic bias [74]. Restricted by the regulation system and liability laws, the integration of human precision with machine efficiency would probably be a future direction of diagnostic medicine, which will enable health care become more prospective, accurate, and available [75, 76].
The digital revolution in medical imaging: The Role of artificial intelligence (AI) in the future of radiology: A subject review · 2026 · DOIFuture research should focus on validating these findings in larger cohorts, exploring Page 14/26 the biological roles of key metabolites, and further optimizing the combined models for clinical application.
Combining radiomics based on high-resolution computed tomography with plasma metabolomics for diagnosing subtypes of rheumatoid arthritis-associated interstitial lung disease · 2026 · DOI1 Limitations This study employed multimodality for AI-assisted prediction of tumor malignancy. A limitation of this study is that all experiments were conducted on a single publicly available dataset.
Multimodal Deep Learning Radiomics Nomogram for Preoperative Breast Cancer Prediction Using Ultrasound Imaging and Clinical Data · 2026 · DOI[6]. Maximum and perpendicular lung- short-axis diameters were measured on axial window images. All measurements were performed on thin-section CT (slice thickness ≤ 1.5 mm) using lung windows (width 1500–1600 Hounsfield units (HU), level −600 to −700 HU). Scanner models and reconstruction settings are summarised in Supplementary Table S1. Semi-automated volumetry was performed with manual refinement to exclude adjacent vessels, bronchi and chest wall structures, consistent with prior volumetric repro- ducibility studies [19]. T1 was defined as the CT closest to surgery for resected cases and the last eligible follow-up scan for surveillance cases. Three-dimensional regions of interest (ROIs) were segmented using ITK-SNAP [20], and nodule volume was calculated as voxel count within the ROI multiplied by voxel volume. Inter-observer agreement was assessed in a double-read subset of 30 nodules for volumetry, mean attenuation, and long- and short-axis diameter measurements at baseline and follow-up. Agreement was excellent across all asses- sed metrics (ICC range, 0.967–0.999; Supplementary Table S2).
CT-based interpretable delta-radiomics model for risk stratification of pulmonary ground-glass nodules: a multicentre study · 2026 · DOIRadiologists, as well as everyone involved with the use of AI in a clinical setting, require familiarity with AI systems to leverage them safely and effectively in clinical practice. AI excels at processing large datasets and handling repetitive tasks, while radiologists provide judgment, context, and adaptability for complex cases, as well as the essential task of data curation. Combining these strengths can lead to more accurate, efficient, and reliable medical imaging practices, accelerating the translation of AI models into clinical practice [12, 13, 83, 90]. Overcoming the limitations of paediatric oncology data The challenges of AI development in paediatric oncology imaging stem from limited data availability and the timeconsuming, labour-intensive process of producing annotations, problems exacerbated by the need for datasets diverse enough to represent the entire paediatric age spectrum. These challenges can be addressed in several ways: increasing the availability of data, applying (AI) methods that generate synthetic data or learn effectively from limited examples, or combining both strategies in a complementary manner. A primary approach to expanding data availability involves multi-institutional collaboration, which can help create datasets that are more diverse across age, demographics, and clinical characteristics. These initiatives rely on standardised protocols and ethical agreements to ensure privacy, reproducibility, and harmonisation. Furthermore, documenting datasets with “datasheets” that detail their composition, collection, and intended uses is critical for responsible implementation. Several dedicated initiatives and platforms now support sharing of paediatric imaging data. Examples of open datasets and online platforms offering downloadable data across a range of disease types, including oncology, are provided in Table 2 and Table 3. These resources can represent a substantial source of labelled and unlabelled data for diverse AI training purposes. Text-based initiatives like the Pediatric Cancer Data Commons, which integrate clinical data across tumour types, are particularly valuable. Their impact could be strengthened by coupling them with dedicated imaging databases, creating larger and more diverse datasets for robust AI development. Such paediatric resources are urgently needed, given that large, dedicated imaging datasets have already become indispensable tools for advancing AI research in adult oncology [64, 89, 100]. Federated infrastructures, furthermore, enable collaborative AI development across institutions without sharing sensitive patient data, preserving privacy while leveraging heterogeneous, decentralised datasets. This decentralised approach allows models to be trained locally within each participating institution, with only aggregated model updates exchanged across sites [9, 12].
Future work will explore advanced multimodal medical image fusion techniques and transformer-based architectures, including multi-level visual transformers for joint analysis of heterogeneous radiographic data, to further enhance diagnostic performance and scalability.
Hybrid diagnostic framework for bone cancer detection using deep learning and radiomics analysis · 2026 · DOIFrontiers in Radiology 14 frontiersin.org Liu et al. 10.3389/fradi.2026.1782678 to optimize the workflow. Third, different scanners, protocols, and patient populations remains to be verified. We will conduct multi-center external validation with standardized imaging harmonization strategies in future work. Second, the retrospective design may introduce selection bias, and the sample size may insufficiently capture the heterogeneity of rare PSOL subtypes. Meanwhile, manual segmentation, despite excellent reproducibility, is labor-intensive for large-scale application; we will explore automated segmentation algorithms in larger multi-center cohorts the biological mechanisms underlying the selected radiomic features require further radiogenomic validation. Additionally, SHAP analysis can only explain the model’s decision logic and feature contributions, but cannot infer direct biological causality between features and PSOL malignant behavior, which warrants further pathological and radiogenomic verification. Finally, we did not perform systematic pathological subtype stratification for benign and malignant PSOLs. Given the significant radiomic and metabolic heterogeneity across different pathological subtypes, we will conduct subtype-specific subanalysis further optimize the model’s robustness and diagnostic performance.
Development and validation of a nomogram using interpretable machine learning to integrate CT radiomics and PET metabolic parameters for predicting benign-malignant differentiation of pulmonary space-occupying lesions · 2026 · DOIThis study has several limitations that should be acknowl- edged. First, as a retrospective analysis, it carries the potential for selection bias, and prospective validation is therefore essential to confirm the clinical utility of our models. Second, the sample size of the independent external validation cohort is relatively modest, which Wang et al. BMC Oral Health (2026) 26:606 may affect the precision of the performance estimates and represents a limitation for broader statistical gener- alization. On the methodological front, several deliberate simplifications were made to balance feasibility with per- formance. Our DLR model analyzed 2D axial slices rather than full 3D volumes, which may not capture the com- plete spatial heterogeneity of tumors. Future studies with larger, prospectively collected cohorts should explore 3D volumetric DL approaches to determine whether they offer incremental predictive value. Similarly, radiomic features were extracted solely from the primary tumor, omitting potential signal from the lymph nodes them- selves. Incorporating perinodal or nodal imaging features could further improve model precision. The data aug- mentation strategy was also conservative, employing only flips to maintain anatomical plausibility. Exploring more advanced techniques may improve robustness. Addition- ally, the models were developed using conventional MRI sequences (T2WI and CET1), and future integration of advanced imaging or genomic data could enhance pre- dictive power. Finally, while our use of a strict training- test split (with feature selection confined to the training partition) provides a robust and unbiased estimate of model generalizability, future studies with larger cohorts could employ nested cross-validation designs to further optimize feature selection stability and model hyperpa- rameters, potentially enhancing reproducibility.
Deep learning radiomics based on multimodal MRI for preoperative prediction of N stage in tongue squamous cell carcinoma: a multicenter study · 2026 · DOIVisual inspection of high attention-weight tiles revealed patterns consistent with genetic aberrations, but interpretation remained challenging; systematic correlation studies linking attention map features to specific genetic mutations (e.g., BRAF, NRAS, GNAQ) in Spitz tumors should be conducted to improve model interpretability.
The prevalence of Spitz tumor subtypes in the dataset differs from general population prevalence due to specialized consultation center caseload; external validation on geographically diverse, community-based dermatopathology cohorts with natural disease prevalence is needed to assess generalizability of the Spitz tumor AI classification model.
The accuracy threshold required to justify AI model implementation costs was not quantitatively determined; simulation studies should systematically investigate the minimum genetic aberration classification accuracy needed for cost-effective deployment of AI-based test selection recommendations in pathology departments.
The simulation experiment demonstrating workflow efficiency gains was limited to Spitz tumors only; extending the AI-based ancillary diagnostic test recommendation approach to other melanocytic tumor subtypes and IHC stains (BRAF, BAP1, β-catenin) should be evaluated to determine applicability across broader dermatopathology practice.
The inclusion criterion of only IHC stain-confirmed or molecularly-analyzed cases introduced selection bias, as conventional melanomas are not routinely genetically characterized in clinical practice; validation of the Spitz tumor AI classification model on unselected, prospectively-collected melanocytic lesions should be conducted to assess real-world performance.
The dataset remains comparatively small despite being the largest Spitz tumor AI study; training on substantially larger datasets should be pursued to evaluate whether increased sample size can improve predictive accuracy for the challenging diagnostic category and genetic aberration classification tasks in Spitz tumor AI models.
The AI model for genetic aberration prediction achieved only 0.55 accuracy; incorporating positional embeddings to capture lesion morphology at lower magnification levels should be systematically evaluated to determine if this architectural modification improves classification performance for Spitz tumor genetic aberration prediction.
The AJCC staging model is constrained to a small number of classification groups and utilizes only TNM variables, limiting how well it can forecast the relative prognoses of specific patients given a range of different predictive variables.
Most-cited papers in Radiomics and Machine Learning in Medical Imaging
- METhodological RadiomICs Score (METRICS): a quality scoring tool for radiomics research endorsed by EuSoMII · Insights into Imaging · 2024 · 263 citations
- Foundation model for cancer imaging biomarkers · Nature Machine Intelligence · 2024 · 172 citations
- Deep learning for cephalometric landmark detection: systematic review and meta-analysis · Clinical Oral Investigations · 2021 · 161 citations
- CSWin-UNet: Transformer UNet with cross-shaped windows for medical image segmentation · Information Fusion · 2024 · 157 citations
- Automated real-world data integration improves cancer outcome prediction · Nature · 2024 · 156 citations
- Deep Learning in Breast Cancer Imaging: State of the Art and Recent Advancements in Early 2024 · Diagnostics · 2024 · 137 citations
- Conditional Diffusion Models for Semantic 3D Brain MRI Synthesis · IEEE Journal of Biomedical and Health Informatics · 2024 · 114 citations
- Rolling-Unet: Revitalizing MLP’s Ability to Efficiently Extract Long-Distance Dependencies for Medical Image Segmentation · Proceedings of the AAAI Conference on Artificial Intelligence · 2024 · 100 citations
- Artificial intelligence-based MRI radiomics and radiogenomics in glioma · Cancer Imaging · 2024 · 87 citations
- Enhancing NSCLC recurrence prediction with PET/CT habitat imaging, ctDNA, and integrative radiogenomics-blood insights · Nature Communications · 2024 · 83 citations
Most recent work
- PySERA: Open-source standardized python library for automated, scalable, and reproducible handcrafted and deep radiomics · Computer Methods and Programs in Biomedicine · 2026
- Artificial Intelligence–Assisted Lung Nodule Evaluation on Low-Dose Chest CT in Asymptomatic Individuals: A Prospective Randomized Controlled Trial · American Journal of Roentgenology · 2026
- A new era of precision diagnosis and treatment for lung cancer: artificial intelligence-driven multimodal data integration and clinical applications · Cell Death & Disease · 2026
- Assessment of Robustness of <scp>MRI</scp> Radiomic Features in Four Abdominal Organs: Impact of Deep Learning Reconstruction and Segmentation · Journal of Magnetic Resonance Imaging · 2026
- Ultrasound-based deep learning radiomics for the differential diagnosis of benign and malignant subpleural pulmonary lesions · Frontiers in Oncology · 2026
- Artificial Intelligence in ALK-Rearranged NSCLC: Forecasting Response and Resistance · Cancers · 2026
- Multimodal Deep Learning Radiomics Nomogram for Preoperative Breast Cancer Prediction Using Ultrasound Imaging and Clinical Data · International Journal of Computational Intelligence Systems · 2026
- Detection of errors in organs at risk delineations for radiotherapy for clinical trial reviews · Physics in Medicine & Biology · 2026
- SPECT and PET imaging of Alzheimer’s disease revisited: from biomarkers to artificial intelligence-based prediction · Annals of Nuclear Medicine · 2026
- Artificial Intelligence in Oral Cancer: A Systematic Review · Journal of Dental Health and Oral Research · 2026
Find a gap in your own Radiomics and Machine Learning in Medical Imaging sub-topic
This page shows what the Radiomics and Machine Learning in Medical Imaging literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →