Open research questions in Psychometric Methodologies and Testing
1,430 unresolved questions extracted from the limitations and future-work sections of 9,462 Psychometric Methodologies and Testing papers in our library. Each links back to the study that raised it.
What the literature leaves open
Discovering the minimum number of experts that are necessary to validate whether an item is essential or unnecessary to be part of a questionnaire, - Investigating whether panel sample members possess appropriate consistency for specific items
Discovering the critical number of respondents to validate an item in a questionnaire: the binomial cut-level content validity proposal · 2026 · DOIThe challenge of making valid inferences from large-scale assessment data, which is not designed for individual-level inferences. The need to account for sampling and imputation error when working with large-scale assessment data. The complexity of working with plausible values, which require careful consideration of the underlying psychometric and score-generation procedures.
Zooming Out On Education: Making Valid Psychological Inferences From Large-Scale Assessment Data · 2026 · DOILarge-scale assessments have limitations due to selection probabilities and nonresponse - Sampling and imputation error need to be considered when analyzing large-scale assessment data - No single guide can be written that is equally applicable across all large-scale assessments
Zooming Out On Education: Making Valid Psychological Inferences From Large-Scale Assessment Data · 2026 · DOIThe study identifies the time-consuming nature of conducting SOMAs, particularly the extraction of statistical results from meta-analyses. The study notes the challenge of ensuring the accuracy of data extraction from meta-analyses. The study highlights the need for practical guidance on the responsible integration of LLMs into data extraction processes.
Automated Data Extraction by Large Language Models: Assessing Accuracy in Comparison to Human Experts Using the Example of Visible Learning · 2026 · DOIThe study only included 156 meta-analyses - 44 studies were excluded due to missing full-text articles - The study relied on a specific dataset (Visible Learning) - The generalizability of the results to other datasets and domains is unknown
Automated Data Extraction by Large Language Models: Assessing Accuracy in Comparison to Human Experts Using the Example of Visible Learning · 2026 · DOIThe study had a small sample size of 105 tenth-grade students, - The high correlations among dimensions and the characteristics of the empirical sample may limit the interpretation of the results, - The study did not test MCAT outside of simulations, - The technical challenge of constructing an MCAT, including complex calculations for ability estimation and item selection, - The need for careful planning and a large amount of test data to create MCAT
Constructing a Computerized Adaptive Test for Multidimensional Mathematical Competence: Development, Simulation, and Validation · 2026 · DOIFurther research is needed to test MCAT outside of simulations, - Future studies should investigate the use of MCAT with larger and more diverse samples, - Research should focus on developing more efficient methods for constructing MCAT, - Studies should examine the effectiveness of MCAT in different educational settings
Constructing a Computerized Adaptive Test for Multidimensional Mathematical Competence: Development, Simulation, and Validation · 2026 · DOIOne challenge identified is the difficulty of developing an overall goodness-of-fit test for item response theory models. Another challenge is the need to evaluate assumptions and model misfit in a practical and substantive way.
Book Review: Review of Item Response Theory: Foundations for Psychologists and Social Scientists (2nd Edition) by Susan E. Embretson and Steven P. Reise (2025) · 2026 · DOIThe gap in the literature is not explicitly stated in the provided text. However, the book appears to address the need for an accessible and practical introduction to item response theory for psychologists and social scientists.
Book Review: Review of Item Response Theory: Foundations for Psychologists and Social Scientists (2nd Edition) by Susan E. Embretson and Steven P. Reise (2025) · 2026 · DOICeiling effects above the 15% threshold for all items. No known-group differences survived correction for multiple testing. The need for a systematic inclusion of children's perceptions within broader environmental-health monitoring activities.
Psychometric properties of the Italian Place Standard Tool for Children: validation in primary school children within the PNC Clima national programme · 2026 · DOIFurther studies should investigate the relationships between environment, climate, and health from a One Health perspective - Future research should include a larger and more diverse sample of children - The study of the Italian PST-C could be extended to other countries and languages
Psychometric properties of the Italian Place Standard Tool for Children: validation in primary school children within the PNC Clima national programme · 2026 · DOIOne challenge is the potential for biased results when using random-intercept cross-lagged panel models (RI-CLPM). Another challenge is the lack of conclusive evidence for the relationship between executive functions and symptoms of psychiatric disorders. A third challenge is the need for careful interpretation of RI-CLPM results to avoid premature conclusions.
Inconclusive effects between executive functions and symptoms of psychiatric disorders in random-intercept cross-lagged panel models: a simulated reanalysis and comment on Halse et al. (2022) · 2025 · DOIResearchers should use triangulation to scrutinize findings from analyses of observational data, - Future studies should carefully interpret RI-CLPM results, - Researchers should consider using latent change score models and models of spurious longitudinal associations
Inconclusive effects between executive functions and symptoms of psychiatric disorders in random-intercept cross-lagged panel models: a simulated reanalysis and comment on Halse et al. (2022) · 2025 · DOIThe impact of varied item wording on participants' choices among specific options remains unexplored. The study aims to address this gap by investigating the effects of item wording on participants' latent target traits and the scale's reliability. The research gap is significant, as the combination of positively and negatively worded items is widely used in Likert scales.
How does item wording affect participants’ responses in Likert scale? Evidence from IRT analysis · 2024 · DOIHowever, the relative behavior of RDIF compared with more flexible residual‐based procedures remains unclear, especially under less ideal conditions such as shorter tests, skewed ability distributions, sample imbalance, difficult item profiles, and nonlinear DIF.
Residual‐Based DIF Detection as a Model‐Agnostic Framework: A Comparison of Parametric, Semi‐Parametric, and Nonparametric Methods · 2026 · DOIHowever, it relies on parameter estimates from sample data and does not account for sampling variability.
Abstract While significant attention has been given to test equating to ensure score comparability, limited research has explored equating methods for rater‐mediated assessments, where human raters inherently introduce error.
IRT Observed‐Score Equating for Rater‐Mediated Assessments Using a Hierarchical Rater Model · 2025 · DOIPrevious research reported conflicting findings on the direction of bias and what contributes to it.
The findings suggest that two items may need revision, and the results must be interpreted considering differences in students’ cultural backgrounds, even though the questionnaires were administered in the same language, which warrants further research.
Validity Study on the Students’ Attitudes Toward Mathematics Scale for English-Speaking Countries · 2025 · DOIAlthough researchers use RMSEA to compare two different models, no studies have compared RMSEA and RDR methods.
Comparison of Factor Retention Methods in Exploratory Factor Analysis: RMSEA, Root Deterioration per Restriction and Parallel Analysis · 2024 · DOIIf little is known about the data a priori, then this statistic should not be used. The tests of these assumptions are not well known and, furthermore, they are probably sensitive to non normality and sample size like their univariate equiv alents.
Future research should investigate the difference between using recall-based rather than recognition-based measures of leader behavior.
Prototypes and scripts: The effects of alternative methods of processing information on rating accuracy · 1987 · DOIThe conclusions made from any Monte Carlo study are necessarily limited by the design of the study.
Improper Solutions in the Analysis of Covariance Structures: Their Interpretability and a Comparison of Alternate Respecifications · 1987 · DOIEven if this were not the case, insufficient information is provided to allow evaluation of these studies.
The challenge of interpreting SAT score trends over time due to changes in exam content. The need for a scalable audit of longitudinal comparability.
Artificial test-takers as transformed controls: measuring SAT difficulty drift and student performance · 2026 · DOI
Most-cited papers in Psychometric Methodologies and Testing
- Evaluating Structural Equation Models with Unobservable Variables and Measurement Error · Journal of Marketing Research · 1981 · 64,623 citations
- Structural equation modeling in practice: A review and recommended two-step approach. · Psychological Bulletin · 1988 · 34,708 citations
- A Coefficient of Agreement for Nominal Scales · Educational and Psychological Measurement · 1960 · 32,341 citations
- A new criterion for assessing discriminant validity in variance-based structural equation modeling · Journal of the Academy of Marketing Science · 2014 · 31,617 citations
- Structural Equation Models with Unobservable Variables and Measurement Error: Algebra and Statistics · Journal of Marketing Research · 1981 · 10,597 citations
- Short screening scales to monitor population prevalences and trends in non-specific psychological distress · Psychological Medicine · 2002 · 8,868 citations
- Random effects structure for confirmatory hypothesis testing: Keep it maximal · Journal of Memory and Language · 2013 · 8,486 citations
- Alternative Ways of Assessing Model Fit · Sociological Methods & Research · 1992 · 7,365 citations
- What is coefficient alpha? An examination of theory and applications. · Journal of Applied Psychology · 1993 · 6,470 citations
- Assessing measurement model quality in PLS-SEM using confirmatory composite analysis · Journal of Business Research · 2019 · 4,594 citations
Most recent work
- On the Consistency of Automatic Scoring with Large Language Models · Educational and Psychological Measurement · 2026
- A Tutorial on Estimating the Precision of Individual Test Scores for Anyone Constructing and Using Psychological Tests · Psychometrika · 2026
- Decoding Item Fit Statistics in Generalized Partial Credit and Graded Response Models · Measurement and Evaluation in Counseling and Development · 2026
- Technology-based versus paper-pencil: sources of mode effects in large-scale assessment · International Journal of Mathematical Education in Science and Technology · 2026
- Using parametric logistic item response theory model trees to detect item parameter heterogeneity in a high-stakes reading comprehension test · International Journal of Testing · 2026
- A tutorial for building multi-type PLS-SEM in social research: validating and testing reflective and formative constructs in a single model · Quality & Quantity · 2026
- Gender differences in willingness to guess revisited: Heterogeneity in a high stakes professional setting · Journal of Economic Behavior & Organization · 2026
- Bad Mood Rising? Assessing Scalar Invariance Violations with Comparative Democratic Support Data · Public Opinion Quarterly · 2026
- Enhancing Psychometric Analysis with Interactive SIA Modules · Psychometrika · 2026
- semfindr: An R Package for Identifying Influential Cases in Structural Equation Modeling · Multivariate Behavioral Research · 2026
Find a gap in your own Psychometric Methodologies and Testing sub-topic
This page shows what the Psychometric Methodologies and Testing literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →