Open research questions in Advanced Statistical Methods and Models
34 unresolved questions extracted from the limitations and future-work sections of 1,439 Advanced Statistical Methods and Models papers in our library. Each links back to the study that raised it.
What the literature leaves open
Future studies are encouraged to expand the dataset by including longer observation periods or multi-country data to improve model generalizability and the stability of hyperparameter optimization. Additional explanatory variables, such as renewable energy consumption, industrial activity, urbanization, carbon pricing, or climate policy indicators, should also be considered to better represent the determinants of CO₂ emissions. Furthermore, future research should compare the present baseline models with more advanced machine learning and deep learning approaches, such as Elastic Net, Partial Least Squares, Random Forest, Gradient Boosting, XGBoost, or Long Short-Term Memory (LSTM), to examine whether higher predictive accuracy can be achieved while maintaining model interpretability.
Evaluation of Linear Regression and Ridge Regression as Baselines for Predicting Indonesia's CO₂ Emissions · 2026 · DOIFuture research should explore the applicability of glmmLasso across diverse educational contexts and consider the development of advanced penalized regression techniques to address challenges such as the optimal selection count thresholds and the appropriate size of ICC within multilevel data structures.
Teachers’ team innovativeness in TALIS 2018: An empirical and simulation study using glmmLasso for multilevel data · 2025 · DOIA Monte Carlo experiment is conducted to evaluate the performance of these estimators and residuals in finite samples on the influence of outliers by considering contaminated data under a perturbation scheme to generate outliers were carried out and confirm that the proposed regression model seems to be a new robust alternative for modeling continuous data limited to the unit interval.
A general and unified parameterization of the beta distribution: A flexible and robust beta regression model · 2025 · DOIAdditionally, analytic relationships between total and level-specific versions of MLM R-squared measures have not been clarified, despite such relationships becoming increasingly important to understand when there are more levels.
The developments in this article are for the DWLS (diagonally weighted least squares) estimator, a popular limited information categorical estimation method.
Hence, assessing or comparing accuracy based on the MSE (which is the mean of squared errors) is insufficient and even inadequate because we should be interested not only in the average but in the whole distribution of prediction errors.
Although there has been no consensus on the best way to construct standardized logistic regression coefficients, there is now sufficient evidence to suggest a single best approach to the construction of a standardized logistic regression coefficient that can be used in the same way across a broad range of problems as the standardized linear regression coefficient and also to suggest the adequacy of other approaches for limited purposes.
When sparse data have to be fitted to a log-linear or latent class model, one cannot use the theoretical chi-square distribution to evaluate model fit, because with sparse data the observed cross-table has too many cells in relation to the number of observations to use a distribution that only holds asymptotically.
For this reason, even the most elaborate statistical treatment of its results m i g h t remain insufficient alike for proper interpretation of complex population phe- nomena, unless one continually attempts to validate them in the light of biological information about the underlying processes.
What happens if some of the assump- tions regarding u are relaxed? What happens, for ex- ample, if we relax the assumption that E(ujuj+k) = 0 if k # O? I n other words, what heppens if we assume that the errors are serially correlated? It is beyond the scope of this paper to go into this problem.
Use of Dummy Variables in Testing for Equality between Sets of Coefficients in Linear Regressions: A Generalization · 1970 · DOIFuture research may consider extensions to high-dimensional settings, dependent data, robustness un- der model misspecification, and generalized semiparametric isotonic models of the form E(Yi | Xi, Zi) = H(g(Xi; θ) + m(Zi)), where H(·) is a known inverse link function.
Residual-based efficient and powerful independence testing in multivariate isotonic semiparametric nonlinear regression · 2026 · DOITo address these scalability limitations, we develop sparse Multivariate Granger Causality (sMVGC), a novel method premised on the assumption that true causal connections between signals are sparse, thereby constraining the candidate search space and improving scalability.
The paper lacks theoretical justification or comparison of why the proposed PCA-Ridge combinations should theoretically outperform existing methods.
Performance of Proposed Ridge – PCA Estimators: Simulation Evidence and Real Data Applications · 2026 · DOIThe simulation results are available on request but for ease of comparison the results are summarized in Table 1, suggesting incomplete presentation of simulation methodology and results.
Performance of Proposed Ridge – PCA Estimators: Simulation Evidence and Real Data Applications · 2026 · DOIEstimated coefficients from the two forms may therefore vary widely, because of their different foci, relative arithmetic versus relative geometric means.
Multiplicative Models For Continuous Dependent Variables: Estimation on Unlogged versus Logged Form · 2017 · DOIIn this work, we show with practical applications that many disparate models, including but not limited to the ones mentioned earlier, can be fitted using gllamm.
For a widely used item response model, when r is small and multidimensional tables are sparse, the proposed statistics have accurate empirical Type I errors, unlike Pearson’s X 2 .
A straightforward interpretation of this phenomenon is lacking, in part due to the unavailability of a closed form for the resulting GEE estimates.
Despite the value of these works, their methods are limited by the required distributional assumptions, by their complexity in implementation, and by the unknown distributions of the estimators.
Structural Equation Models That are Nonlinear in Latent Variables: A Least-Squares Estimator · 1995 · DOIThis model comparison is insufficient for model evaluation: In large samples virtually any model tends to be rejected as inadequate, and in small samples various competing models, if evaluated, might be equally acceptable.
She has managed, by careful structuring, to give the reader ready access to a wide range of studies, guiding him discreetly through a welter of often conflicting findings.
Most-cited papers in Advanced Statistical Methods and Models
- Significance tests and goodness of fit in the analysis of covariance structures. · Psychological Bulletin · 1980 · 13,503 citations
- Recent Advances in Quantile Regression Models: A Practical Guideline for Empirical Research · The Journal of Human Resources · 1998 · 918 citations
- Extracting the Variance Inflation Factor and Other Multicollinearity Diagnostics from Typical Regression Results · Basic and Applied Social Psychology · 2017 · 916 citations
- Doubly robust difference-in-differences estimators · Journal of Econometrics · 2020 · 895 citations
- Performance of the Modified Poisson Regression Approach for Estimating Relative Risks From Clustered Prospective Data · American Journal of Epidemiology · 2011 · 425 citations
- Inference in Multiscale Geographically Weighted Regression · Geographical Analysis · 2019 · 376 citations
- Some Contributions to Efficient Statistics in Structural Models: Specification and Estimation of Moment Structures · Psychometrika · 1983 · 309 citations
- Cluster-robust inference: A guide to empirical practice · Journal of Econometrics · 2022 · 288 citations
- Revisiting Gaussian copulas to handle endogenous regressors · Journal of the Academy of Marketing Science · 2021 · 258 citations
- Limited Information Goodness-of-fit Testing in Multidimensional Contingency Tables · Psychometrika · 2006 · 257 citations
Most recent work
- Statistical Power, Reliability, and Model Assumptions in Multiple Linear Regression: Assessing and Addressing Common Issues · Measurement and Evaluation in Counseling and Development · 2026
- Simultaneous heterogeneity and reduced-rank learning for multivariate response regression · Journal of Multivariate Analysis · 2026
- Scalable and Robust Regression Models for Continuous Proportional Data · Journal of the American Statistical Association · 2026
- On the definition of median for a random interval · Information Sciences · 2026
- Performance of Proposed Ridge – PCA Estimators: Simulation Evidence and Real Data Applications · International Journal of Development Mathematics (IJDM) · 2026
- Beyond MSE in Poisson Ridge Regression: New Ridge Parameter Estimators with Additional Distributional Performance Criteria · Mathematics · 2026
- Improved Data-Driven Shrinkage Estimators for Regression Models Under Severe Multicollinearity · Mathematics · 2026
- Navigating Multicollinearity in Linear Regression Models: Implications for Big Data Analysis · Contemporary Mathematics · 2026
- A method enabling computation of linear rates of change of spatial averages on visual field patterns that have varying test locations over time · medRxiv · 2026
- Handling missing data, skewness, and outliers in medical research: A robust factor analysis approach using the canonical fundamental skew-t distribution · Statistical Methods in Medical Research · 2026
Find a gap in your own Advanced Statistical Methods and Models sub-topic
This page shows what the Advanced Statistical Methods and Models literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →