Validation gaps in Computer Science
433 open validation research questions in Computer Science — gaps in reproducing, validating, or independently confirming findings — extracted from 295 papers in our local library. Below are representative open questions, each linked to the paper that raised it.
Representative open questions
Showing 30 of 433 — one per source paper, highest-quality first.
- Federated learning for privacy-preserving skin cancer classification using deep neural networks (2026) · doi
The model's performance on the ISIC dataset is evaluated using various metrics (accuracy, AUC, F1-score), but the robustness of these metrics to different conditions (e.g., varying image quality) is not assessed.
- Hybrid Deep Model for Pain Intensity Classification Using Fused ECG, EMG, and GSR Signals (2026) · doi
The study evaluates pain intensity classification using only 5-fold cross-validation on a single proprietary dataset without reporting dataset size, participant demographics, pain induction methods, or signal sampling rates. External validation on publicly available pain-related physiological signal datasets (e.g., BioVid, UNBC-McMaster) is necessary to assess generalization of the hybrid model across different pain assessment protocols and populations.
- Turbulence closure in Reynolds-averaged Navier–Stokes and flow inference around a cylinder using physics-informed neural networks and sparse experimental data (2026) · doi
The Reynolds-force model was trained and validated exclusively on cylinder flow data at Re ≈ 300−300,000 with the near-wall region remaining laminar; the generalization of the trained neural network closure to separated flows around bluff bodies with different geometries or highly turbulent near-wall regions has not been demonstrated.
- FPGA-Enabled Machine Learning Applications in Earth Observation: A Systematic Review (2026) · doi
Most CNN-based FPGA implementations for remote sensing do not benchmark against prior FPGA-based CNN architectures, making it impossible to identify commonly compared network designs and establish reproducible performance baselines. Standardized benchmarking protocols comparing different CNN models and FPGA devices for Earth observation tasks are needed to enable meaningful cross-study comparisons.
- Securing IoT Devices with PUFs: Mitigating Aging and Tampering through Cryptography and Machine Learning (2026) · doi
The condition N minπi ≥ 5 for χ² approximation validity is stated, but no empirical testing is provided to determine whether this conservative condition remains sufficient for IoT PUF sequences with reduced entropy due to aging-induced bit-flip rates or tampering-induced environmental stress.
- Blockchain-integrated machine learning framework for transparent smart contract vulnerability detection (2026) · doi
The SmartBugs-Wild dataset evaluation revealed highly skewed cluster distributions (35,499 contracts in Cluster 1 vs. 149 in Cluster 2) with only 34.15% variance explained by the first two principal components, yet the impact of this structural heterogeneity on model generalization across clusters has not been evaluated. Cross-cluster validation performance between models trained on different structural archetypes of smart contracts should be investigated.
- Can deep learning-based segmentation and classification improve the detection of renal cortical abnormalities? (2026) · doi
Although Grad-CAM visualizations confirm the DenseNet205 model focuses on cortical scarring regions, the paper does not compare the localization accuracy of model activations against radiologist-annotated scar boundaries or validate whether the model's attention patterns align with clinically significant scarring thresholds. Quantitative analysis of Grad-CAM activation overlap with expert scar delineations is needed for clinical acceptance.
- Inferring High-Dimensional Dynamic Networks Changing with Multiple Covariates (2026) · doi
The paper identifies that TP53 exhibits unclear connection patterns with several genes (ATG4A, RAD9B, STX6, TRIM13, XRCC4, CES2, SESN2, FBXO22, XRCC6, GRB2, PRKAA1, TAF3, NOX4) across cancer groups and radiation doses, but provides no mechanistic validation of these radiation dose-dependent network rewiring patterns through experimental confirmation or functional genomics approaches.
- COGNITIVE INFRASTRUCTURE AND THE RECURSIVE TRANSFORMATION OF KNOWLEDGE COMMUNICATION: GENERATIVE AI IN SCIENTIFIC PUBLISHING (2026) · doi
The paper argues that certification mechanisms are exposed as inadequate when generative systems enter writing and evaluation, triggering an arms-race dynamic where detection tools stimulate further innovation in generative capability. However, it does not characterize the specific failure modes of current integrity detection methods against evolving AI text generation techniques or provide benchmarks for measuring detection evasion rates across different manuscript types and disciplinary norms.
- Prediction of sedimentation concentration profiles in inclined suspension systems: A data-driven neural network framework (2026) · doi
The ANN model was trained and tested exclusively within a single experimental domain (glycerin–water 92% v/v, glass microspheres 212–800 µm, 20% v/v solids), meaning high accuracy reflects interpolation rather than extrapolation; generalization to other rheologies, particle morphologies, volumetric concentrations, or field-scale conditions in inclined suspension systems remains untested.
- Large language model based machine translation for universal multilingual understanding and translation quality enhancement (2026) · doi
The paper evaluates machine translation quality using multiple metrics (BLEU, COMET, BERTScore) but does not systematically analyze which metrics best correlate with human judgment for different language pairs or establish metric-specific performance trends for universal multilingual understanding and translation quality assessment.
- Leveraging machine learning to enhance aerosol classification using Single-Particle Mass Spectrometry (2026) · doi
While the dataset contains 18,827 labeled spectra across 20 aerosol types, generalization performance of the machine learning framework on SPMS measurements from different atmospheric environments, seasons, or geographic locations has not been evaluated. Cross-dataset validation is needed to assess whether the supervised and semi-supervised models maintain classification fidelity for soot and feldspars across diverse real-world deployment scenarios.
- Quantum Information Framework for Neural Network Generalization: A Comprehensive Experimental Analysis (2026) · doi
Training is conducted with a single optimizer (Adam) and fixed hyperparameter ranges (learning_rate=0.01, weight_decay=0.0); the effect of different optimization algorithms and regularization intensities on quantum information metrics and their correlation with generalization remains untested.
- AI in Cybersecurity: A Systematic Review and Conceptual Audit Model (2026) · doi
The model claims to balance AI-driven technology adoption with adequate cybersecurity safeguards, yet provides no validation dataset, benchmark, or test cases demonstrating how organizations with different maturity levels, infrastructure complexity, or resource constraints should implement the Anti-Sheriff framework operationally.
- Large Language Models for Combinatorial Optimization: A Systematic Review (2026) · doi
Prompt learning using metaheuristics (reference [167]) requires evaluation on larger combinatorial optimization benchmark sets and comparison against traditional prompt engineering techniques to establish when metaheuristic-driven prompt optimization outperforms baseline LLM configurations.
- Explainable machine learning for tracking spatial variation in leaf chlorophyll fluorescence within temperate deciduous forest canopies (2026) · doi
The Random Forest and XGBoost models for predicting chlorophyll fluorescence parameters were trained exclusively on temperate deciduous forest datasets. These models require retraining and validation across diverse ecosystems (croplands, grasslands, wetlands) to assess transferability and determine whether spectral reflectance-ChlF relationships remain consistent across different plant functional types and canopy architectures.
- Machine-learning-based reconstruction of Ming-dynasty defensive corridors in Yuxian (2026) · doi
The paper validates defense corridors against 181-232 military defense sites using buffer coverage analysis (72-83% coverage at 2-5 km), but does not evaluate model performance using other validation metrics (precision, recall, F1-score), cross-validation across temporal periods (early vs. late Ming), or compare predictions against archaeologically-excavated versus documented-only heritage sites.
- Comparative analysis of deep learning algorithms for rolling element bearing fault classification under variable loads and speeds (2026) · doi
The robustness analysis used purely additive white Gaussian noise at 1 dB, 3 dB, and 5 dB SNR levels, but this does not represent the full range of mechanically-induced disturbances in industrial bearing systems such as load torque ripple, rotational speed fluctuation, structural resonance, or shaft misalignment. Future work should evaluate deep learning models for rolling element bearing fault classification under these realistic non-stationary mechanical phenomena with frequency drift, harmonic amplification, and impulsive components.
- The adoption of artificial intelligence methods in entrepreneurship research: current state and pathways forward (2026) · doi
The cascading term substitution problem in VOSviewer—where excluding multi-word terms causes shorter term variants to capture occurrences of excluded terms (e.g., 'artificial neural network' → 'neural network' → 'network')—is identified but lacks systematic evaluation of how this substitution cascade distorts co-occurrence networks and clustering in entrepreneurship AI literature reviews.
- A two-stage deep learning model for risk identification in green supply chain finance (2026) · doi
The GAN-SAE synthetic data augmentation component is applied to the training set, but the paper does not report how synthetic data quality or distribution fidelity affects model robustness when applied to out-of-sample industries in the 2021-2024 external validation; ablation studies isolating GAN contribution versus SAE feature extraction are needed for green supply chain finance applications.
- LLM-Powered Silent Bug Fuzzing in Deep Learning Libraries via Versatile and Controlled Bug Transfer (2026) · doi
TransFuzz's oracle validation achieves 71.42% precision due to three specific failure modes: (1) oracle design errors causing logical flaws in bug classification, (2) code migration failures that fail to preserve bug characteristics across different PyTorch APIs, and (3) insufficient API documentation leading to LLM misjudgment. These sources of false positives (28.58% rate) require targeted improvements in oracle construction and API information enrichment for silent bug detection in deep learning libraries.
- Quantum-SpinalNet: a hybrid deep learning approach for mammographic breast cancer detection (2026) · doi
The attention map validation mentioned in section 5.4 is not detailed in the excerpt. The paper should specify how attention maps from the Q-SpinalNet were clinically validated against radiologist annotations to ensure that highlighted regions correspond to actual tumor boundaries and suspicious features rather than spurious artifact activations.
- Ocean: Object-aware Anchor-free Tracking with Matching-relation Learning (2026) · doi
While Ocean++ is adapted to pixel-level tracking on VOT-2020 by grafting a segmentation network branch, the integration strategy and potential conflicts between the matching-relation learning mechanism and pixel-level segmentation requirements are not discussed or evaluated.
- Automated design of heuristics for resource-constrained project scheduling problem via regression algorithms (2026) · doi
The CPU time comparison between regression-based heuristics (executed on standard desktop) and the genetic algorithm (executed on HPC system with parallel computing) is acknowledged as non-comparable. A controlled computational benchmark under identical hardware and parallel processing conditions is needed to definitively establish the computational efficiency claims of regression-based heuristics versus GA-based methods for large resource-constrained project scheduling.
- From unstructured text to structured reasoning: a hybrid knowledge graph for Indonesian sentencing analysis (2026) · doi
The relevance filtering stage uses SMOTE oversampling to achieve 100% recall and 99.30% precision, but the paper does not evaluate whether this extreme precision-recall trade-off introduces cascading errors in downstream entity extraction when applied to court decisions with novel or unusual sentence structures not represented in the training set.
- Neural Network Tools in the Arsenal of a University Teacher (2026) · doi
The paper states that AI tool effectiveness depends on user competency but does not empirically investigate or measure how different levels of user expertise in AI-assisted literature search (query formulation, result validation, critical evaluation) affect the quality outcomes when using neural network-based research tools in academic settings.
- Securing Fog-assisted IoT: An Adaptable and Efficient Threat Identification Approach (2026) · doi
The assumption that fog servers process traffic immediately without transmission or queuing delays is unrealistic in production environments. The DEL models require re-evaluation under actual network conditions with variable queuing delays, packet loss rates, and bandwidth constraints typical in real fog-assisted IoT deployments.
- A Robust Hybrid Deep Learning Model for Multiclass Depression Classification from Speech Audio (2026) · doi
Statistical significance testing and confidence interval estimation were not performed due to single train–test split and limited dataset size; rigorous cross-validation and statistical validation protocols must be implemented to establish confidence in the multiclass depression classification results.
- On the interface between linguistics, computer science and psychiatry: analyzing textual key-factors affecting BERT-based classification of schizophrenia in social media texts (2026) · doi
The interaction between text length, discourse genre, and schizophrenia linguistic markers appears additive rather than interactive, but this relationship has not been formally tested across controlled genre conditions with minimum-length thresholds. Systematic manipulation of both text length and topic/genre type is needed to establish genre-specific minimum-length requirements for reliable BERT-based classification.
- Artificial Intelligence (AI) Based Multi-Layered Approaches for Privacy Preservation in Federated Learning (2026) · doi
The ablation study demonstrates synergistic effects when combining federated learning, differential privacy, and homomorphic encryption, but does not investigate how the relative contributions of each component vary across different dataset sizes, feature dimensionality, or number of federated clients in the multi-layered privacy framework.
Working on one of these gaps? Review it with us.
Science AI Journal reviews manuscripts in one pass with 8 specialised AI agents calibrated on 69,000+ real peer reviews.