Decision Sciences · Research topic

Open research questions in Psychometric Methodologies and Testing

1,430 unresolved questions extracted from the limitations and future-work sections of 9,462 Psychometric Methodologies and Testing papers in our library. Each links back to the study that raised it.

What the literature leaves open

  • Discovering the minimum number of experts that are necessary to validate whether an item is essential or unnecessary to be part of a questionnaire, - Investigating whether panel sample members possess appropriate consistency for specific items

    Discovering the critical number of respondents to validate an item in a questionnaire: the binomial cut-level content validity proposal · 2026 · DOI
  • The challenge of making valid inferences from large-scale assessment data, which is not designed for individual-level inferences. The need to account for sampling and imputation error when working with large-scale assessment data. The complexity of working with plausible values, which require careful consideration of the underlying psychometric and score-generation procedures.

    Zooming Out On Education: Making Valid Psychological Inferences From Large-Scale Assessment Data · 2026 · DOI
  • Large-scale assessments have limitations due to selection probabilities and nonresponse - Sampling and imputation error need to be considered when analyzing large-scale assessment data - No single guide can be written that is equally applicable across all large-scale assessments

    Zooming Out On Education: Making Valid Psychological Inferences From Large-Scale Assessment Data · 2026 · DOI
  • The study identifies the time-consuming nature of conducting SOMAs, particularly the extraction of statistical results from meta-analyses. The study notes the challenge of ensuring the accuracy of data extraction from meta-analyses. The study highlights the need for practical guidance on the responsible integration of LLMs into data extraction processes.

    Automated Data Extraction by Large Language Models: Assessing Accuracy in Comparison to Human Experts Using the Example of Visible Learning · 2026 · DOI
  • The study only included 156 meta-analyses - 44 studies were excluded due to missing full-text articles - The study relied on a specific dataset (Visible Learning) - The generalizability of the results to other datasets and domains is unknown

    Automated Data Extraction by Large Language Models: Assessing Accuracy in Comparison to Human Experts Using the Example of Visible Learning · 2026 · DOI
  • The study had a small sample size of 105 tenth-grade students, - The high correlations among dimensions and the characteristics of the empirical sample may limit the interpretation of the results, - The study did not test MCAT outside of simulations, - The technical challenge of constructing an MCAT, including complex calculations for ability estimation and item selection, - The need for careful planning and a large amount of test data to create MCAT

    Constructing a Computerized Adaptive Test for Multidimensional Mathematical Competence: Development, Simulation, and Validation · 2026 · DOI
  • Further research is needed to test MCAT outside of simulations, - Future studies should investigate the use of MCAT with larger and more diverse samples, - Research should focus on developing more efficient methods for constructing MCAT, - Studies should examine the effectiveness of MCAT in different educational settings

    Constructing a Computerized Adaptive Test for Multidimensional Mathematical Competence: Development, Simulation, and Validation · 2026 · DOI
  • One challenge identified is the difficulty of developing an overall goodness-of-fit test for item response theory models. Another challenge is the need to evaluate assumptions and model misfit in a practical and substantive way.

    Book Review: Review of Item Response Theory: Foundations for Psychologists and Social Scientists (2nd Edition) by Susan E. Embretson and Steven P. Reise (2025) · 2026 · DOI
  • The gap in the literature is not explicitly stated in the provided text. However, the book appears to address the need for an accessible and practical introduction to item response theory for psychologists and social scientists.

    Book Review: Review of Item Response Theory: Foundations for Psychologists and Social Scientists (2nd Edition) by Susan E. Embretson and Steven P. Reise (2025) · 2026 · DOI
  • Ceiling effects above the 15% threshold for all items. No known-group differences survived correction for multiple testing. The need for a systematic inclusion of children's perceptions within broader environmental-health monitoring activities.

    Psychometric properties of the Italian Place Standard Tool for Children: validation in primary school children within the PNC Clima national programme · 2026 · DOI
  • Further studies should investigate the relationships between environment, climate, and health from a One Health perspective - Future research should include a larger and more diverse sample of children - The study of the Italian PST-C could be extended to other countries and languages

    Psychometric properties of the Italian Place Standard Tool for Children: validation in primary school children within the PNC Clima national programme · 2026 · DOI
  • One challenge is the potential for biased results when using random-intercept cross-lagged panel models (RI-CLPM). Another challenge is the lack of conclusive evidence for the relationship between executive functions and symptoms of psychiatric disorders. A third challenge is the need for careful interpretation of RI-CLPM results to avoid premature conclusions.

    Inconclusive effects between executive functions and symptoms of psychiatric disorders in random-intercept cross-lagged panel models: a simulated reanalysis and comment on Halse et al. (2022) · 2025 · DOI
  • Researchers should use triangulation to scrutinize findings from analyses of observational data, - Future studies should carefully interpret RI-CLPM results, - Researchers should consider using latent change score models and models of spurious longitudinal associations

    Inconclusive effects between executive functions and symptoms of psychiatric disorders in random-intercept cross-lagged panel models: a simulated reanalysis and comment on Halse et al. (2022) · 2025 · DOI
  • The impact of varied item wording on participants' choices among specific options remains unexplored. The study aims to address this gap by investigating the effects of item wording on participants' latent target traits and the scale's reliability. The research gap is significant, as the combination of positively and negatively worded items is widely used in Likert scales.

    How does item wording affect participants’ responses in Likert scale? Evidence from IRT analysis · 2024 · DOI
  • However, the relative behavior of RDIF compared with more flexible residual‐based procedures remains unclear, especially under less ideal conditions such as shorter tests, skewed ability distributions, sample imbalance, difficult item profiles, and nonlinear DIF.

    Residual‐Based DIF Detection as a Model‐Agnostic Framework: A Comparison of Parametric, Semi‐Parametric, and Nonparametric Methods · 2026 · DOI
  • However, it relies on parameter estimates from sample data and does not account for sampling variability.

    Added Value of Subscores: Can We Accurately Evaluate It? · 2026 · DOI
  • Abstract While significant attention has been given to test equating to ensure score comparability, limited research has explored equating methods for rater‐mediated assessments, where human raters inherently introduce error.

    IRT Observed‐Score Equating for Rater‐Mediated Assessments Using a Hierarchical Rater Model · 2025 · DOI
  • Previous research reported conflicting findings on the direction of bias and what contributes to it.

    From Item Estimates to Test Operations: The Cascading Effect of Rapid Guessing · 2025 · DOI
  • The findings suggest that two items may need revision, and the results must be interpreted considering differences in students’ cultural backgrounds, even though the questionnaires were administered in the same language, which warrants further research.

    Validity Study on the Students’ Attitudes Toward Mathematics Scale for English-Speaking Countries · 2025 · DOI
  • Although researchers use RMSEA to compare two different models, no studies have compared RMSEA and RDR methods.

    Comparison of Factor Retention Methods in Exploratory Factor Analysis: RMSEA, Root Deterioration per Restriction and Parallel Analysis · 2024 · DOI
  • If little is known about the data a priori, then this statistic should not be used. The tests of these assumptions are not well known and, furthermore, they are probably sensitive to non normality and sample size like their univariate equiv alents.

    The Analysis of Repeated Measures Designs Involving Multiple Dependent Variables · 1987 · DOI
  • Future research should investigate the difference between using recall-based rather than recognition-based measures of leader behavior.

    Prototypes and scripts: The effects of alternative methods of processing information on rating accuracy · 1987 · DOI
  • The conclusions made from any Monte Carlo study are necessarily limited by the design of the study.

    Improper Solutions in the Analysis of Covariance Structures: Their Interpretability and a Comparison of Alternate Respecifications · 1987 · DOI
  • Even if this were not the case, insufficient information is provided to allow evaluation of these studies.

    Wide Range Achievement Test (WRAT‐R), 1984 Edition · 1986 · DOI
  • The challenge of interpreting SAT score trends over time due to changes in exam content. The need for a scalable audit of longitudinal comparability.

    Artificial test-takers as transformed controls: measuring SAT difficulty drift and student performance · 2026 · DOI

Most-cited papers in Psychometric Methodologies and Testing

Most recent work

Find a gap in your own Psychometric Methodologies and Testing sub-topic

This page shows what the Psychometric Methodologies and Testing literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.

Open the Research Gap Finder →

Related topics in Decision Sciences

1,430 open questions have been extracted from the limitations and future-work passages of 9,462 Psychometric Methodologies and Testing papers in our library. Each one below links back to the study that raised it, so you can read the original claim in context.

Tools for your next paper

Compare the category — Honest roundups of the AI research tools, ours listed alongside the alternatives.

Command palette

Jump anywhere, run any action.