Open research questions in Meta-analysis and systematic reviews
133 unresolved questions extracted from the limitations and future-work sections of 2,480 Meta-analysis and systematic reviews papers in our library. Each links back to the study that raised it.
What the literature leaves open
Future research could explore a wider range of persuasive devices, including those requiring contextual or longitudinal anal- ysis, and use qualitative methods (e. Future studies might address this by counterbalancing ques- tion order and including a broader range of persuasive strategies, from ethically neutral to problematic, to better distinguish concerns specific to QRPs from general attitudes toward persuasion. Future research should address these limitations by including larger, multidisciplinary samples and using designs with a higher ecological validity Bajraktari & Munafò (2026). While our analyses did not account for these depen- dencies, we acknowledge that this may have influenced standard error estimates.
Persuasion in Scientific Journal Articles: Exploring Persuasiveness and Informativeness among Authors and Readers · 2026 · DOIHowever, choosing between available NMR models is complex, as each model addresses a similar, but unique research question, and the performance of available NMR models under varying network structures, between-study heterogeneity, and interaction assumptions remains unclear.
Assessing the Impact of Model Assumptions in Network Meta-Regression: A Simulation Study · 2026This assumption helps explain why the traditional three-level model, which does not account for between-study variability in the moderator effect, produced inflated Type I error rates. While the extent to which within-study moderator effects vary across studies is not well established, it is reasonable to expect such variability.
Meta-regression with categorical moderators and dependent effect sizes: A simulation study · 2026 · DOIEmpirical evidence supports Johnston and Pennypacker’s (1993) assertion that researchers’ biases may affect how they design, conduct, and interpret their SCED studies (Tincani et al., 2025; Travers et al., 2025; Mahoney, 1977; Shadish et al., 2016). When SCED researchers explicitly acknowledge where a study falls on the inductive-deductive research continuum, this serves to transparently disclose the fundamental aims of the study, enhancing the validity and believability of reported findings. In confirmatory and deductive studies, public disclosure of researcher pre- dictions in the form of directional research questions and hypotheses functions as a guardrail to ensure transparent research and reporting practices. Recent methodologi- cal guidelines have recognized the importance of distinguishing between inductive and deductive studies in SCED research (Ledford et al., 2023; Tincani et al., 2024); however, recognition of this distinction is not yet ubiquitous in SCED (e.g., Tate et al., 2016; What Works Clearninghouse, 2023). Accordingly, future SCED research guidelines should more explicitly encourage authors to describe whether a study is deductive, inductive, or a combination of both. The research examples presented in this commentary illustrate that the inductive – deductive research continuum is not a dichotomous concept in that studies are either one or the other, as many SCED studies include varying elements of induction and deduction. For example, researchers might conduct a deductive study to confirm whether an established intervention improves some socially important behavior. If the intervention does not yield optimal experimental or clinical effects under cer- tain conditions (i.e., the researcher encounters a boundary of the intervention), the researcher might introduce a novel variation of the procedure to improve therapeutic responding in the study (e.g., Crozier & Tincani, 2005), an element of induction. Most critical is that researchers transparently disclose their procedures and decision- making processes to permit unbiased interpretation of the findings. This transparency can be achieved by publishing a study preregistration prior to data collection and updating it as changes occur (Tincani et al., 2024), and/or by clearly documenting any study modifications and their rationales in the final research report (see Tincani et al., 2025). Finally, Johnston and Pennypacker (1993) cautioned that directional research questions and hypotheses leading to “yes” or “no” answers could constrain more nuanced and rich interpretations of functional relations in SCED experiments. This concern has considerable merit and the guidance provided in this commentary should not be interpreted to mean that the results of deductive and confirmatory studies should only be interpreted as responses to “yes” or “no” questions. In many cases, Journal of Behavioral Education1 3 ambiguous inconsistent, or otherwise variable patterns in responding may not permit definitive “yes” or “no” answers, or they may highlight undiscovered environment- behavior relations to be further explored. Researchers conducting confirmatory and deductive studies are encouraged to fully interpret their results as permitted by a nuanced visual analysis, a hallmark of SCED. Acknowledgements This work was supported by a grant from the U.S. Department of Education, ED- GRANTS-040821-001, Office of Special Education Programs, under priority CFDA 84.325D. Author contributions M.T. conceived the idea for the manuscript, conducted the literature search, and drafted and edited the work. Data Availability No datasets were generated or analysed during the current study.
We want to end with some recommendations for applied researchers. 1. Investigate the variation of the effect across cases and studies. Besides inves- tigating the overall treatment effect, it is important to study to what extent the effect can be generalized across cases and studies. Using prediction intervals can help to get an insight in interpreting the between-case and between-study vari- ances. Studying moderator effects of case and study characteristics can help to understand better for whom and under what conditions a treatment works, and the proportion of variance explained gives an idea of how well we succeed in this. 2. Standardize data only if necessary. If across studies data are measured on a com- parable and intuitive scale, we recommend analyzing the unstandardized data, because standardizing can lead to flawed parameter estimates. 3. Use Hedges’ correction factor when analyzing standardized data. For standard- as an ized data, the meta-analysis of effect size estimates using σ 2 = n0+n1 n0.n1 β 1 estimate of the sampling variance will give very similar results as the meta- analysis of raw data. Hedges’ correction factor will remove not only the bias in (cid:31) Journal of Behavioral Education1 3 the treatment effect estimates, but also to improve the estimates of the variance components. 4. Be prudent in using variance estimates, as they can be biased, even when using Hedges’ correction factor. For the scenarios that were simulated (i.e., for con- tinuous data without time trends and autocorrelation) we found that the variance estimates are especially biased when there are for each participant 10 or fewer observations per phase. For other scenarios (e.g., for discrete data or data with time trends or autocorrelation) we cannot give strong recommendations without further research. Yet, also for these scenarios we recommend researchers to be prudent, especially when the number of observations is small. 5. Consider using a parametric bootstrap procedure. For the scenarios that were simulated, we found that the bias in variance estimates can be largely removed by using a parametric bootstrap. For other conditions, we suggest using a para- metric bootstrap as a sensitivity analysis: caution is especially warranted if very different variance estimates are obtained, whereas confidence in the estimates is increased if similar variance estimates are obtained. Supplementary Information The online version contains supplementary material available at h t t p s : / / d o i . o r g / 1 0 . 1 0 0 7 / s 1 0 8 6 4 - 0 2 6 - 0 9 6 1 9 - w . Acknowledgements The computational resources and services used in this work were provided by the VSC (Flemish Supercomputer Center), funded by the Research Foundation-Flanders (FWO) and the Flemish Government-department EWI. Author Contributions WVdN & MM conceptualized the study; WVdN wrote the simulation code, did the analyses and wrote a first version of the manuscript; MM critically reviewed the codes, analyses and manuscript. Data Availability No datasets were generated or analysed during the current study.
It may also explin why we only found 12 reporitng guidelines had undergone qualita- tive user testing, and why these qualitative components were often small, perhaps limited to a single question like “Please add your comments and suggestions in the free text below” [48] or “Any other feedback?” [52]. This may help explain why a recent audit of report- ing guidelines in the EQUATOR Network database found that few undergo piloting at all and that pilots are rarely reported in detail [13].
What influences whether researchers adhere to healthcare reporting guidelines successfully? A thematic synthesis · 2026 · DOIFuture studies could consider using ‘in the moment’ methods like think-aloud tasks. Firstly, influences 6 and 8–12 were not explored by any non-qualitative questions. We have already described how our review is limited by the availability of literature and the relative thinness of some studies’ qualitative analysis. Our Chinese database searches yielded no relevant studies, and we found no studies on this topic published in Spanish or Portuguese.
What influences whether researchers adhere to healthcare reporting guidelines successfully? A thematic synthesis · 2026 · DOIFindings suggest there is insufficient empirical evidence to make credible causal inferences about the effects of MR on policy-relevant outcomes, highlighting significant research gaps that impede our understanding of whether this policy is functioning as intended.
RCQRS may be beneficial in any of the following circumstances: (a) the timeframe to complete the synthesis is short; (b) the scope of inquiry is not fully defined, (c) the availability of suitable literature is beyond the screening capacity of the reviewers; or (d) the availability of literature is sparse and reviewers seek to extrapolate insights from other areas.
Reverse chronology quota record screening for realist synthesis: Fostering causally rich extrapolations with a diverse and contemporaneous sample of literature · 2026 · DOIThese conclusions are limited to the presented bench- mark of classical TF-IDF–based supervised baselines and a zero-shot open-weight LLM, both evaluated with mini- mal parameter tuning.
Comparing supervised machine learning and large language models in title-abstract screening · 2026 · DOIThe results of this study are only partially comparable to those of other studies, as many of these vary in terms of Aigner et al. Systematic Reviews (2026) 15:190 Page 8 of 14 a b Fig. 3 Estimator scores on average and per dataset. a Recall and specificity averaged across all datasets. In comparison with a human benchmark, all models achieve similar recall, and logistic regression, random forest, and support vector machine achieve similar specificity. Narrower confidence intervals indicate how consistent a model performed. b Recall and specificity by dataset. Scores are higher in the pancreatic surgery, animal depression, and ADHD dataset compared to the other three. With supervised machine learning, recall and specificity tend to be closer together than with Llama, which consistently achieves higher recalls, even on the lower performing datasets. The boxplots have been produced by the matplotlib.axes.Axes.boxplot function in python on 1000 bootstrap samples of model predictions per dataset. The box extends from the first quartile (Q1) to the third quartile (Q3) with the center line representing the median. The whiskers extend from the box to the farthest data point lying within the 1.5 × the inter-quartile range (IQR) between Q1 and Q3. Outliers have been present but omitted from the figure for clarity datasets, models and other methods that they used. We intentionally chose datasets with varying class imbal- ances compared to the 3–6% common in health-related SRs [59] to see whether different inclusion rates would influence classification performance. was employed for the LLM evaluation. Further research could include more elaborate prompting approaches, such as few-shot examples or in-context learning.
Comparing supervised machine learning and large language models in title-abstract screening · 2026 · DOIFuture research should focus on curating datasets, such as SYNERGY, for reuse in studies on the automation of title/abstract-screening. A follow-up study could use this study’s results to investigate how attributes in the data- sets and inclusion criteria contribute to classification per- formance. The dataset on animal depression is especially relevant in this context because all models in this study and in other studies have achieved classification perfor- mance comparable to that of expert human reviewers using this dataset. Especially for living SRs, future studies should train on earlier update cycles and evaluate on later waves in order to test temporal generalization under real- istic prospective conditions.
Comparing supervised machine learning and large language models in title-abstract screening · 2026 · DOIMISSING COVARIATES IN META-REGRESSION 40 Although we presented the feasibility of adapting the MB-MI approach for meta-regression and discussed the scenarios where the MB-MI approach could be more beneficial or perform at least as well as agnostic MI, at least for now, its benefits need to be further examined under additional focal models and data structures.
Exploring the Use of Multiple Imputation for Handling Missing Covariates in Meta-regression with Dependent Effect Sizes · 2026 · DOIFifth, while this study did test prompts against MA/SRs across four different medical fields (psychiatry, endocrinology, pulmonology, orthodontics), future research should examine LLM screening performance against more MA/SRs within and beyond medicine.
Prompt engineering of large language models for paper screening in medical meta-analyses and systematic reviews: A prospective comparative study · 2026 · DOI— — — • Discuss any limitations related to the living mode • Describe any changes since the preceding version to the implications of the results for practice, policy, and future research • Describe and justify any planned changes to review methods in upcoming review versions • Indicate and justify whether the LSR is being retired from the living mode following the publication of the current version, if applicable — — — • Describe the sources of financial or non-financial support and the roles of funders or sponsors in each of the versions of the LSR • Describe the competing interests of review authors and how they were managed for each of the versions of the LSR • Describe any changes since the preceding version to the accessibility of data, code…
Extension of the PRISMA 2020 statement for living systematic reviews (PRISMA-LSR): checklist and explanation · 2024 · DOIClarifications are warranted throughout the Interim Cochrane RRMG guidance to ensure that users with various experience levels can understand and apply its recommendations accordingly.
Evaluation of the interim Cochrane rapid review methods guidance—A mixed‐methods study on the understanding of and adherence to the guidance · 2023 · DOIThe results suggest that when starting with a PROSPERO entry and where no trials have been screened for inclusion, automated methods can reduce workload, but additional processes are still needed to efficiently identify trial registrations or trial articles that meet the inclusion criteria of a systematic review.
A comparison of machine learning methods to find clinical trials for inclusion in new systematic reviews from their <scp>PROSPERO</scp> registrations prior to searching and screening · 2023 · DOITo avoid the condition-dependent, all-or-none choice between competing methods and conflicting results, we extend robust Bayesian meta-analysis and model-average across two prominent approaches of adjusting for publication bias: (1) selection models of p-values and (2) models adjusting for small-study effects.
Robust Bayesian meta‐analysis: Model‐averaging across complementary publication bias adjustment methods · 2022 · DOIThe aim of this study was to analyze publicly available justifications for stabilizing a Cochrane review, with the ultimate goal of helping to make decisions about whether the update of any SR is warranted.
How to decide whether a systematic review is stable and not in need of updating: Analysis of Cochrane reviews · 2020 · DOIAlthough most teams were confident their selected targets would provide useful information to the clinician, not one recommendation was similar: both the number (0-16) and nature of selected targets varied widely.
Time to get personal? The impact of researchers choices on the selection of treatment targets using the experience sampling methodology · 2020 · DOIThe paper acknowledges the role of online tools in shaping nuclear medicine compliance with EBM methods, but does not specify what features, accessibility levels, or integration workflows these tools should contain to support the conduct and publication of EBM-compliant nuclear medicine research.
The excerpt mentions governmental incentives driving adherence to EBM standards and AUCs in nuclear medicine, but does not identify which specific regulatory mechanisms, reimbursement structures, or policy interventions would most effectively incentivize EBM adoption among practitioners and institutions.
The paper proposes that journals will eventually reject articles that do not demonstrate EBM conformance, but does not detail the specific standardized criteria, checklists, or reporting guidelines (beyond STARD, CONSORT, SQUIRE) that nuclear medicine journals should uniformly adopt to enforce this quality threshold.
While the paper advocates for professional societies to promote EBM through Appropriate Use Criteria (AUCs) and practice guidelines, it does not specify how to systematically evaluate the impact of these AUCs on nuclear medicine clinical decision-making or measure compliance rates across different healthcare settings and institutions.
The paper identifies that many current nuclear medicine practitioners completed their formal education before EBM innovations were introduced, but does not specify what targeted continuing education curricula or competency assessments are needed to ensure practicing physicians achieve proficiency in EBM methodologies such as ROC analysis, diagnostic accuracy study design, and systematic review appraisal.
Most-cited papers in Meta-analysis and systematic reviews
- <i>PRISMA2020</i> : An R package and Shiny app for producing PRISMA 2020‐compliant flow diagrams, with interactivity for optimised digital transparency and Open Synthesis · Campbell Systematic Reviews · 2022 · 2,861 citations
- A scoping review of scoping reviews: advancing the approach and enhancing the consistency · Research Synthesis Methods · 2014 · 2,495 citations
- Outlier and influence diagnostics for meta-analysis · Research Synthesis Methods · 2010 · 2,469 citations
- Effects of Computerized Clinical Decision Support Systems on Practitioner Performance and Patient Outcomes · JAMA · 2005 · 2,246 citations
- Scientific procedures and rationales for systematic literature reviews (SPAR‐4‐SLR) · International Journal of Consumer Studies · 2021 · 1,882 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resources · Research Synthesis Methods · 2019 · 1,822 citations
- Empirical Evidence for Selective Reporting of Outcomes in Randomized Trials · JAMA · 2004 · 1,391 citations
- Sample size, power and effect size revisited: simplified and practical approaches in pre-clinical, clinical and laboratory studies · Biochemia Medica · 2020 · 1,304 citations
- Practical Clinical Trials · JAMA · 2003 · 1,227 citations
- The Mass Production of Redundant, Misleading, and Conflicted Systematic Reviews and Meta‐analyses · Milbank Quarterly · 2016 · 1,183 citations
Most recent work
- Prompt engineering of large language models for paper screening in medical meta-analyses and systematic reviews: A prospective comparative study · Research Synthesis Methods · 2026
- Assessing the properties of the prediction interval in random-effects meta-analysis · Research Synthesis Methods · 2026
- An international consensus on core reproducibility items in research · PLOS Biology · 2026
- How comprehensive is comprehensive? Prevalence of searches beyond peer-reviewed journal articles and bibliographic sources: a cross-sectional study of systematic reviews · Research Synthesis Methods · 2026
- Concordance between target trial emulation and randomised controlled trials: systematic review and meta-analysis · BMJ · 2026
- Letter Regarding “Ten Reasons Why Prospective Randomized Studies in Surgery Are Flawed and Fundamentally Different From Drug Trials” · The Journal Of Hand Surgery · 2026
- A philosopher asks whether therapy should be evaluated like drugs · The British Journal of Psychiatry · 2026
- The publication symmetry test: a simple editorial heuristic to combat publication bias · Journal of Clinical and Translational Research · 2026
- Giving credit where credit’s due - recognition of patient partners in health research · Research Involvement and Engagement · 2026
- Scoping Reviews in Health Professions Education · The Clinical Teacher · 2026
Find a gap in your own Meta-analysis and systematic reviews sub-topic
This page shows what the Meta-analysis and systematic reviews literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →