Open research questions in Statistical Methods and Inference
99 unresolved questions extracted from the limitations and future-work sections of 1,038 Statistical Methods and Inference papers in our library. Each links back to the study that raised it.
What the literature leaves open
The varying nature of time-dependent covariates poses complications in assessing their effect on survival outcomes. Existing methods require parametric or semiparametric modeling of the relationship between the time-dependent covariate and the survival outcome.
Differing endpoints for each individual. Limited number of time points and genes in the dataset. Difficulty in applying traditional multivariate analysis techniques.
The paper identifies the challenge of designing statistically efficient functional estimates while remaining agnostic to the structure of the sampling distribution. The authors highlight the challenge of understanding the trade-offs between different debiasing methods. The paper discusses the challenge of developing more robust and efficient debiasing methods.
The maximum likelihood estimate may not exist in high-dimensional settings. The estimator may have high aggregate bias. The paper needs to provide strong empirical evidence for the effectiveness of the rescaled estimator.
Jeffreys’ Prior Penalty for High-Dimensional Logistic Regression: A Conjecture About Aggregate Bias · 2026 · DOICausal inference problems are challenging due to the need to estimate the average treatment effect. Density ratio estimation is a key component of causal inference. Riesz regression is a general tool for Riesz representer estimation.
The conventional Gaussian quasi-likelihood function is sensitive to contamination from discontinuous variations. There is a need for robustified versions of the conventional Gaussian quasi-maximum-likelihood estimator.
To extend the strategy of augmenting observed data with optimization decision variables beyond linear program. To explore the application of the approach to other types of optimization problems, such as semi-definite program.
Existing methods are inadequate for handling combinatorial response data. The lack of a link function that connects a linear or functional predictor with a probability respecting the combinatorial constraints.
The literature lacks model averaging procedures for either the response mechanism or the conditional quantile function. The existing methods can hardly skirt the estimation of MNAR mechanism.
Semiparametric model averaging for high-dimensional quantile regression with nonignorable nonresponse · 2026 · DOIThe limitation of in-memory regression schemes in Big Data scenarios. The need for a scalable regression framework that can efficiently compute various regression models.
Model misspecification and selection effect. Lack of finite-sample validity. Inferior interval forecast accuracy at longer forecast horizons.
Conformal prediction for functional time series: application to age-specific mortality rates · 2026 · DOITo study the properties of the proposed estimators in more detail. To apply the proposed estimators to real-world data.
However, the ordinary maximum likelihood estimator (MLE) may suffer from finite-sample bias when few studies are available or the mean event probability is close to the boundary.
Penalized likelihood inference for beta-binomial meta-analysis of proportions of rare events · 2026Together, these results extend the DE-constrained regression paradigm to a broad class of physically motivated linear models and provide a practical estimation tool for fire science and other application areas where mechanistic knowledge is available but data are sparse.
Local Quasi-Linear Models: Kernel Differential Equation Regression and Fire Data Analysis · 2026We then apply the framework to the firebrand burning-rate experiment of Albini (1979), modelling the density-loss curve of wind-driven firebrands with a physically motivated forced-convection ODE; the DE-constrained estimator outperforms local linear regression in the sparsest species-diameter groups, where physical structure is most valuable in compensating for scarce data.
Local Quasi-Linear Models: Kernel Differential Equation Regression and Fire Data Analysis · 2026These identification problems originate from the poorly defined mapping between a structural model and reduced-form parameters.
This study provides an auditable and transferable forecasting protocol for operational researchers and university administrators, helping institutions determine when context-aware forecasting adds practical value under limited data and structural instability.
Forecasting commencing enrolments under data sparsity: a zero-shot time series foundation models framework for higher education planning · 2026 · DOIWe conclude that the field is converging toward integrated frameworks that jointly address dimension reduction, sparsity, dependence, and data imperfections, and that the most consequential open problems lie precisely at the intersections of the four themes.
QUANTILE REGRESSION: A REVIEW OF METHODOLOGICAL ADVANCES IN LINEAR, BAYESIAN, PANEL DATA, AND MEASUREMENT ERROR MODELS · 2026 · DOIThe paper identifies the need for nonparametric estimation of statistical error. The paper discusses the limitations of traditional methods for estimating statistical error.
There is a gap in the application of the knowledge of bias to the prediction context.
The observed inverse relationship between network size (decreasing node count) and network density (increasing interconnectivity) as radiation dose increases from 0 to 2 Gy is descriptive rather than mechanistically explained; the biological reasons why higher radiation selectively eliminates nodes while strengthening remaining gene-gene interactions in cancer versus control groups are not addressed.
The Hamming distance analysis reveals that radiation doses of 0 Gy and 0.05 Gy produce nearly identical network structures with only marginal differences attributed to random variation, but no threshold or statistical test is established to distinguish true dose-dependent network changes from stochastic fluctuations in CVN estimation.
To evaluate the variable selection performance under weak hierarchical structure or no hierarchy structure. To apply the proposed method to other large-scale registries where event times are recorded on a discrete scale.
Gradient boosting-based discrete failure time model for selecting time-varying effects and interactions · 2026 · DOIClassical discrete failure time models often assume proportional hazards and linear additive effects of covariates. The existing methods do not efficiently identify both time-varying and time-independent effects, while preserving interpretability.
Gradient boosting-based discrete failure time model for selecting time-varying effects and interactions · 2026 · DOIThe paper focuses on a binary treatment scenario. The method may not be applicable to multiple treatment options. The approach requires further extension to longitudinal studies and survival data.
Most-cited papers in Statistical Methods and Inference
- Structural Vector Autoregressions: Theory of Identification and Algorithms for Inference · The Review of Economic Studies · 2009 · 727 citations
- African migration: trends, patterns, drivers · Comparative Migration Studies · 2016 · 304 citations
- Anatomy of the Selection Problem · The Journal of Human Resources · 1989 · 295 citations
- Model Averaging and Its Use in Economics · Journal of Economic Literature · 2020 · 265 citations
- Model averaging prediction by K-fold cross-validation · Journal of Econometrics · 2022 · 262 citations
- Response Surface Regressions for Critical Value Bounds and Approximate p‐values in Equilibrium Correction Models1 · Oxford Bulletin of Economics and Statistics · 2020 · 244 citations
- Targeting predictors in random forest regression · International Journal of Forecasting · 2022 · 154 citations
- A Tutorial on Estimating Time-Varying Vector Autoregressive Models · Multivariate Behavioral Research · 2020 · 146 citations
- Revisiting the Gelman–Rubin Diagnostic · Statistical Science · 2021 · 116 citations
- Smoothed quantile regression with large-scale inference · Journal of Econometrics · 2021 · 114 citations
Most recent work
- On multiplicative bias correction in kernel density estimation · Open Research Online (The Open University) · 2026
- Random-weighting bootstrap correction for Granger causality tests in VAR models with time-varying variance of unknown form · Economics Letters · 2026
- CONFIDENCE INTERVALS FOR MULTIPLE CHANGE POINTS IN LINEAR MODELS WITH HETEROSCEDASTIC ERRORS · Econometric Theory · 2026
- Inferring High-Dimensional Dynamic Networks Changing with Multiple Covariates · Applied Mathematics and Statistics · 2026
- Stochastic multiple imputation for latent heterogeneity in interval-censored data: A prior-stratified approach (PS-SMI) · Hacettepe Journal of Mathematics and Statistics · 2026
- Gradient boosting-based discrete failure time model for selecting time-varying effects and interactions · Lifetime Data Analysis · 2026
- Successive classification learning for estimating quantile optimal treatment regimes · Journal of the American Statistical Association · 2026
- A Generalized Discontinuous Hamilton Monte Carlo for Transdimensional Sampling · Journal of Scientific Computing · 2026
- Comparing variable selection and model averaging methods for logistic regression · Proceedings of the National Academy of Sciences · 2026
- Statistical inference for Gaussian Whittle–Matérn fields on metric graphs · Journal of the Royal Statistical Society Series B: Statistical Methodology · 2026
Find a gap in your own Statistical Methods and Inference sub-topic
This page shows what the Statistical Methods and Inference literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →