Open research questions in Statistical Methods and Bayesian Inference
48 unresolved questions extracted from the limitations and future-work sections of 909 Statistical Methods and Bayesian Inference papers in our library. Each links back to the study that raised it.
What the literature leaves open
Missing data and confounding are common in real-world statistical applications, yet few studies have examined how imputation methods perform under time-varying confounding in binary variables, or how missingness mechanism, missing rate, missingness location and sample size jointly affect performance and the underlying identifiability conditions.
Comparing Missing Data Methods for Estimating Average Treatment Effects Under Time-Varying Confounding: A Simulation Study · 2026In the near term, we plan to extend the framework in several directions. First, the risk-targeted calibration layer will be developed and integrated with a fully for- mulated CRC/LTT-based risk control mechanism. This will ensure the creation of a tighter link between empirical measures of downstream performance and sample risk targeted calibration. Second, the framework will incorporate more expressive counterfactual value models. Third, it will be extended from short-horizon glucose alarming to a larger class of medical time-series decision tasks, including other physiological monitoring settings, multimodal patient streams, and a variety of clin- ically structured missing patterns. Future versions of the framework will include more complex delayed confirmations. A more practical approach for deployment would likely involve separating the full EVSI computation from the wearable sensor. Wearable devices typically have limited processing capabilities. They are designed primarily to sense physiological signals and transmit these data to a secondary processing system. Examples include a user’s smartphone, a home gateway, a hospi- tal edge server, or even a cloud-based monitoring service. This type of architecture is most appropriate when resources are constrained in terms of power consumption. Each wearable device can remain focused on its primary functions of sensing and transmitting. The EVSI-based decision process, which requires greater computational capability, is handled by the more powerful secondary processing system. 220 International Journal of Online and Biomedical Engineering (iJOE) iJOE | Vol. 22 No. 7 (2026) Decision Framework Focused on Missing-Data Alarms in Healthcare While delayed confirmation was simulated using fixed time intervals in the current prototype, measurement availability in real clinical settings may vary due to sensor contact loss, patient movement, network instability during data transfer, or workflow-related delays. The EVSI framework can be generalized to account for this variability by taking an expectation over a waiting-time distribution. Instead of assuming a single deterministic delay, the decision algorithm would evaluate the expected value of waiting across possible arrival times of the confirmatory measurement. A more extensive validation process for the framework will be performed using larger healthcare datasets. 7 ACKNOWLEDGMENTS This publication was made possible with the financial support of AKKSHI. Its content is the responsibility of the author, and the opinion expressed therein is not necessarily the opinion of AKKSHI. 8 REFERENCES [1] W. Du, W. Côté, and Y. Liu, “SAITS: Self-attention-based imputation for time series,” Expert Systems with Applications, vol. 219, p. 119619, 2023. https://doi.org/10.1016/ j.
The increasing availability of large-scale epidemiological studies, such as the UK Biobank, poses important modelling challenges, including mixed data types, high dimensionality, and structured missingness.
The reasons why Random Forest performed differently in real data compared to simulated data, and how practitioners should select between RF, ARIMA, and LSTM based on data characteristics, requires further research.
Analysis of Imputation Methods for Missing Not at Random (MNAR) Data: A Comparative Study of Air Pollution Data in Bangkok, Thailand · 2026 · DOIThe methodology tested imputation at missing rates up to 70%, but whether KNN maintains effectiveness at even higher missing rates (>70%) remains unexplored.
Analysis of Imputation Methods for Missing Not at Random (MNAR) Data: A Comparative Study of Air Pollution Data in Bangkok, Thailand · 2026 · DOIKNN showed poor forecasting accuracy for variables characterized by simpler MAR or MCAR patterns, but the underlying reasons for this performance differential are not thoroughly explained.
Analysis of Imputation Methods for Missing Not at Random (MNAR) Data: A Comparative Study of Air Pollution Data in Bangkok, Thailand · 2026 · DOIAlthough much research has considered applications of MI in hierarchical data, little is known about its use in cross-classified data, in which observations are clustered in multiple higher-level units simultaneously (e.
Handling Missing Data in Cross-Classified Multilevel Analyses: An Evaluation of Different Multiple Imputation Approaches · 2023 · DOIIn the cases where simple heuristics are insufficient, the Shake-and-Bake technique outperforms the modified IPFP in terms of time complexity and the accuracy of searching the space of feasible solutions.
Even when one begins with the same random number seed, conflicting findings can be obtained from the same data under an identical imputation model between SAS® and SPSS®.
An Examination of Discrepancies in Multiple Imputation Procedures Between SAS® and SPSS® · 2018 · DOISince the education variable in SIAB is reported for statistical reasons only, it suffers from frequent inconsistent reports and a high and increasing share of missing values.
Reducing the Need for Heuristic Rules – An Iterative Algorithm for Imputing the Education Variable in SIAB · 2015 · DOIPlanned missing designs are becoming increasingly popular, but because there is no consensus on how to implement them in longitudinal research, we simulated longitudinal data to distinguish between strategies of assigning items to forms and of assigning forms to participants across measurement occasions.
Optimal assignment methods in three-form planned missing data designs for longitudinal panel studies · 2014 · DOILittle is known about the performance of PMM in imputing non‐normal semicontinuous data (skewed data with a point mass at a certain value and otherwise continuously distributed).
The final method, expectation maximization (EM), produces asymptotically unbiased estimates, but EM's implementation in MVA is limited to point estimates (without standard errors) of means, variances, and covariances.
Results show that under commonly encountered conditions, a test of fit based on the limited information in the second-order marginals has a Type II error rate that is no higher than the error rate found for full-information test statistics, and that the test statistic given in this paper does not suffer from ill effects of sparseness in the joint frequencies.
3. A Goodness-of-Fit Test for the Latent Class Model When Expected Frequencies are Small · 1999 · DOIAbstract Bayesian analysis provides a robust way to incorporate prior knowledge into statistical models, but eliciting diverse subjective priors for parameters in the unit interval [0, 1] remains lacking.
Missing data is a fundamental challenge in space biology, where high experimental costs, limited sample availability, and tissue allocation constraints produce datasets that are sparse, multimodal, and heterogeneous.
A systematic imputation framework for sparse, multimodal space biology datasets: application to retinal imaging and omics from the RR9 mission · 2026 · DOIThe discrepancy between simulation and real-world results likely reflects the greater complexity, noise, and inter-variable dependence present in real-world data compared with the controlled simulated environment, but this is not fully investigated.
Analysis of Imputation Methods for Missing Not at Random (MNAR) Data: A Comparative Study of Air Pollution Data in Bangkok, Thailand · 2026 · DOIOur results show that while some caution must be exercised when using the Bayesian beta-binomial in meta-analyses with extremely sparse data, the use of a weakly informative prior for the effect parameter is beneficial in terms of mean bias, mean squared error, and coverage.
However, researchers also need ratings that adhere to psychometric standards, such as a certain degree of reliability, and psychometric work with planned missing designs is currently lacking in the literature.
Planning Missing Data Designs for Human Ratings in Creativity Research: A Practical Guide · 2023 · DOIKitagawa-Blinder-Oaxaca decompositions provide consistent estimates also for smaller samples but require assumptions for model specification and, when common support is lacking, for model-based extrapolation.
Comparing the Incomparable? Issues of Lacking Common Support, Functional-Form Misspecification, and Insufficient Sample Size in Decompositions · 2023 · DOIThis approach may enable social scientists to draw new conclusions from sparse data sets with a large number of features, for example, historical or archival sources, online surveys with high attrition rates, or data sets created from Web scraping, which confound traditional imputation techniques.
Sparse Data Reconstruction, Missing Value and Multiple Imputation through Matrix Factorization · 2022 · DOIWhile several simulation studies exist that compare various so-called factor retention criteria under different data conditions, little is known about the impact of missing data on this process.
This approach is occasionally employed in biosciences like plant breeding, but, ironically, has not been established in behavioral sciences despite the close historical connection with factor analysis in these fields.
This sort of single-imputation method has been criticized for producing biased results in other areas of clinical research, but has not been evaluated within the context of alcohol clinical trials, and many alcohol researchers continue to use the missing = heavy drinking assumption.
Methodologists have developed mediation analysis techniques for a broad range of substantive applications, yet methods for estimating mediating mechanisms with missing data have been understudied.
Most-cited papers in Statistical Methods and Bayesian Inference
- When can group level clustering be ignored? Multilevel models versus single-level models with sparse data · Journal of Epidemiology & Community Health · 2008 · 271 citations
- Multiple Imputation With Large Data Sets: A Case Study of the Children's Mental Health Initiative · American Journal of Epidemiology · 2009 · 215 citations
- Best practices for addressing missing data through multiple imputation · Infant and Child Development · 2023 · 158 citations
- Missing Data in Alcohol Clinical Trials: A Comparison of Methods · Alcoholism Clinical and Experimental Research · 2013 · 147 citations
- Predictive mean matching imputation of semicontinuous variables · Statistica Neerlandica · 2014 · 146 citations
- Addressing Data Sparseness in Contextual Population Research · Sociological Methods & Research · 2007 · 135 citations
- A Comparison of Three Popular Methods for Handling Missing Data: Complete-Case Analysis, Inverse Probability Weighting, and Multiple Imputation · Sociological Methods & Research · 2022 · 91 citations
- Dasymetric Modeling and Uncertainty · Annals of the Association of American Geographers · 2013 · 83 citations
- Evaluating the Performances of Missing Data Handling Methods in Ability Estimation From Sparse Data · Educational and Psychological Measurement · 2020 · 74 citations
- Biases in SPSS 12.0 Missing Value Analysis · The American Statistician · 2004 · 74 citations
Most recent work
- Multiple Imputation of Missing Data in Moderated Factor Analysis · Multivariate Behavioral Research · 2026
- Analysis of Linked Files: A Missing Data Perspective · Statistical Science · 2026
- Analysis of Imputation Methods for Missing Not at Random (MNAR) Data: A Comparative Study of Air Pollution Data in Bangkok, Thailand · Journal of Current Science and Technology · 2026
- Bayesian Estimation of Coefficient Alpha Using a Normal Posterior: Non‑normal Distributions · Journal of Modern Applied Statistical Methods · 2026
- Adapting tree-based multiple imputation methods for multilevel data? A simulation study · Behavior Research Methods · 2026
- From Poisson observations to fitted negative binomial distribution · Statistical Papers · 2026
- A stochastic method to estimate a zero-inflated two-part mixed model for human microbiome data · Statistical Modelling · 2026
- A Beta-Binomial Model for Estimating Zero- or One-inflated Pain Trajectories · bioRxiv · 2026
- Methodological Considerations for Quantile Aggregation in Alzheimer Disease Trials · JAMA Neurology · 2026
- A Novel Fusion for Uncertain Dynamic Count Data: Neutrosophic Negative Binomial Mixed-Effects and INAR(p) Models · Journal of the Indian Society for Probability and Statistics · 2026
Find a gap in your own Statistical Methods and Bayesian Inference sub-topic
This page shows what the Statistical Methods and Bayesian Inference literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →