Open research questions in Psychometric Methodologies and Testing
339 unresolved questions extracted from the limitations and future-work sections of 9,068 Psychometric Methodologies and Testing papers in our library. Each links back to the study that raised it.
What the literature leaves open
Third, the possible biases in AI-generated items, which are not yet apparent in the first psychometric examination, should be further studied through more extensive research, through the qualitative analysis of the item content and linguistic nature.
Human vs. AI: Assessing Scale Development for Perceived Risks of ChatGPT in Academic Settings · 2026 · DOIThird, the unexpected equal difficulty level of ID-if-pattern, extend, and abstract items highlights the need for further research on the difficulty of pattern tasks, with different response formats, as part of careful consideration of assessment design. Second, ID-Unit and complete the pattern tasks were not included in Study 2’s version of the assessment, and our assessment did not include a duplicate the pattern task, so conclusions should be made with caution, and future research including all patterning tasks is warranted.
Finally, although the EIRM employed in this study provides a robust analytical framework for examining the interplay between individual learner characteristics and item properties, further research is needed to investigate additional factors that may contribute to mathematics literacy outcomes, including socio-economic status, instructional quality, and other contextual influences.
Items, gender, response time, and styles in play: Exploring student-item interactions in mathematics literacy · 2026 · DOIFuture research could examine the viability of bifactor or multidimensional models to further probe whether more fine-grained facets of L2 buoyancy can be distinguished.
Developing and validating the Second Language Buoyancy Scale (L2BS): Evidence from psychometric and machine learning analyses · 2026 · DOIFuture research could explore in more detail how differentiating between rapid guessing and idling behaviors impacts ability estimates in different contexts. Hence, the results of this study should be replicated with other assessment data. One limitation is that the dataset used to create the six cases consisted entirely of short items, each intended to be completed within ten seconds. Given this short time limit, the RT distributions may not be representative of those from other types of assessments, such as those with a test-level time limit or those where individual items require a longer time to complete.
Disengagement matters: A response-time informed approach to scoring low-stakes assessments · 2026 · DOIIt is recommended that implementation of the proposed framework follow a phased deployment strategy, each stage of deployment must involve three key checks: locally adjusting risk score thresholds, tracking missed cases in high-risk groups, and analyz- ing performance across different subgroups (like gender, language, or location) to ensure fairness. If unfair differences are found, fixes like adjusting thresholds or improving data quality should be made before expanding the system further, beginning with pilot testing in a limited number of schools to assess feasibility, contextual fit, and educator usabil- ity. This should be complemented by targeted teacher and school psychologist training focused on interpreting feature importance, decision paths, and explainability outputs to align model insights with classroom observations [15]. At the policy level, collaboration with educational authorities is essential to establish ethical guidelines, data protection protocols, and structured referral pathways for learners identified as at risk. Future research should prioritize external validation using datasets from diverse regions, academic cohorts, and linguistic contexts [72] to strengthen generalizability. Longitudinal studies are also recommended to assess the framework’s capacity to predict emerging learning risks over time. Finally, periodic model retraining and integration of adaptive learning analytics will be critical to maintaining accuracy, interpretability, and relevance as educational contexts and learner profiles evolve.
A 3-tier machine learning framework for early detection of learning difficulties in basic school settings · 2026 · DOIMoreover, the simulation studies shed new insights into the frequentist behavior of the two DIF indices under conditions that had not been previously explored but which are applicable to many LSA.
Differential item functioning analysis in large scale assessments: a case study for DIF in SABER 11 · 2026 · DOI, 2009) established this factor as rater-specific and stable, but its latent-level relation to dedicated desirable-responding subscales has not been tested.
The Devil Is in the Similarities: A Latent Variable Approach to Detecting and Controlling Halo Bias · 2026 · DOIfor indices limitations that should be considered when interpreting the findings. First, the analysis employed a hybrid measurement specification, combining observed single-indicator variables with reflective multi-item constructs. While this approach is methodologically defensible, such measurement choices may affect construct precision and the stability of estimated relationships. Second, global model fit the estimated model indicated weak overall fit, which constrains the strength of inference and supports interpreting the results cautiously as predictive and than confirmatory. Third, the sample size (N = 71) and the use of a single institutional context limit statistical power and external reduce generalisability. These sensitivity to detect indirect effects, particularly in mediation testing. Accordingly, conclusions regarding unsupported mediated mechanisms should be understood as reflecting insufficient evidence within this sample, rather than definitive absence of mediation. Future studies would benefit from larger, more diverse samples, refined measurement structures, and stronger model fit to enable more robust testing of complex relational and mediated processes. characteristics may also exploratory rather The small sample (N = 71) limits statistical power and increases the likelihood of Type II error. As a result, the absence of statistical significance for some paths should not be interpreted as evidence of no relationship. The findings theory-consistent, should be warranting replication with ideally, longitudinal or experimental designs to support stronger causal inference.
Algorithm reasoning as a learning pathway: a PLS-SEM model linking computational understanding to conceptual explanation in South African mathematics · 2026 · DOIThe effectiveness of integrated instructional strategies (embedding reading comprehension into Mathematics, quantitative reasoning into language classes, working memory drills, metacognitive regulation training) has not been empirically tested through controlled trials to measure their impact on cross-subject performance.
Multivariate Interdependencies of Chinese, Mathematics, and English Performance among High-School Students · 2026 · DOIMachine learning and big data analytics have not been applied to predict individual learning trajectories and enable real-time adaptive teaching that targets the latent academic proficiency factor across Chinese, Mathematics, and English simultaneously within the New Gaokao framework.
Multivariate Interdependencies of Chinese, Mathematics, and English Performance among High-School Students · 2026 · DOIThe generalizability of the single-factor model to other high-stakes assessment systems remains untested; cross-cultural comparisons with systems such as the International Baccalaureate, A-Levels, or AP examinations are required to validate whether domain-general competencies show similar interdependencies across different assessment frameworks.
Multivariate Interdependencies of Chinese, Mathematics, and English Performance among High-School Students · 2026 · DOIMultilevel structural equation modeling must extend beyond individual student data to incorporate school-level factors (instructional quality, teacher training in integrated curricula) and district-level factors (resource allocation, New Gaokao policy implementation) to determine their moderation of the latent academic proficiency factor.
Multivariate Interdependencies of Chinese, Mathematics, and English Performance among High-School Students · 2026 · DOIMultimodal assessment designs integrating cognitive measures (executive function tests), affective measures (motivation scales), and neurophysiological measures (EEG, fMRI) must be developed to disentangle the specific mechanisms underpinning cross-domain transfer among the three Gaokao subjects.
Multivariate Interdependencies of Chinese, Mathematics, and English Performance among High-School Students · 2026 · DOIThe observational longitudinal design precludes causal inference; randomized experiments embedding targeted metacognitive training and adaptive learning interventions are needed to clarify whether domain-general competencies directly drive cross-subject performance interdependencies in Chinese, Mathematics, and English.
Multivariate Interdependencies of Chinese, Mathematics, and English Performance among High-School Students · 2026 · DOIOnly academic performance scores were incorporated without measuring cognitive variables (working memory, metacognition, executive control), affective variables (motivation, self-efficacy), or socioeconomic factors that could mediate cross-domain relationships among the three Gaokao subjects.
Multivariate Interdependencies of Chinese, Mathematics, and English Performance among High-School Students · 2026 · DOIThe study analyzed data from only a single high school in one province; generalizability across different regions, school systems, and provinces within China remains unvalidated for the identified latent academic proficiency factor across Chinese, Mathematics, and English performance.
Multivariate Interdependencies of Chinese, Mathematics, and English Performance among High-School Students · 2026 · DOIThe sliding scale method for identifying low-frequency distractors from high item difficulty questions, as presented in Raymond, Stevens, and Bucak (2019), must be formally incorporated into the integrated methodology to address the approximately 20% of questions per exam with fewer than optimal effective distractors.
A three-category classification system for item fit (infit and outfit) did not prove useful for the presented dataset, requiring validation of whether a competing two-category classification ('acceptable' and 'poor') is more effective for identifying problematic items in multiple choice mathematics assessments.
The utility of standard deviation-based bounds (1 and 2 standard deviations) for item difficulty classification in the integrated methodology requires further investigation to determine if these thresholds effectively predict item classification across varying educational assessment datasets.
The limitations of this study lie in the limited number of participants and the focus on the Analysis phase, which has not yet empirically tested the system's effectiveness. Future research is recommended to advance to the next stage of the ADDIE model by developing a functional prototype of an AI-IRT-based Web CAT system, conducting large-scale system testing in various school settings, and evaluating its impact on student learning outcomes, learning motivation, and assessment fairness.
The Analysis Needs Mapping for the Development of an Artificial Intelligence-Item Response Theory (AI-IRT)-Based CAT Web System for Adaptive Science Assessment · 2026 · DOITherefore, further research is needed to involve a wider range of participants, develop a system prototype, and conduct empirical trials to measure the actual impact of using AI- IRT-based Web CAT on junior high school science learning outcomes. The limitations of this study lie in the limited number of participants and the focus on the Analysis phase, which has not yet empirically tested the system's effectiveness.
The Analysis Needs Mapping for the Development of an Artificial Intelligence-Item Response Theory (AI-IRT)-Based CAT Web System for Adaptive Science Assessment · 2026 · DOIAlthough the CULI test showed sound psychometric functioning, the Rasch results indicated the need for further investigation into certain correct and incorrect choices and the difficulty levels of writing items, which could provide valuable insights for optimising test quality.
Argument-based validation of Chulalongkorn University Language Institute (CULI) test: a Rasch-based evidence investigation · 2025 · DOIDespite the widespread application of BKT in predicting student performance and modeling knowledge acquisition, methods to evaluate the consistency of measurement by BKT models remain elusive.
Since our Rasch analyses were sample-independent, we conclude that the SPIRES survey is a four construct survey and may be treated as such across educational contexts without further need to validate using factor analysis, which could continuously produce inconsistent results.
A Rasch Analysis of the Self-determination, Purpose, Identity, and Engagement in Science (SPIRES) Survey: Instrument Validation and Recommendations · 2025 · DOI
Most-cited papers in Psychometric Methodologies and Testing
- CB-SEM vs PLS-SEM methods for research in social sciences and technology forecasting · Technological Forecasting and Social Change · 2021 · 1,999 citations
- Latent Class Analysis: A Guide to Best Practice · Journal of Black Psychology · 2020 · 1,866 citations
- Applying Bifactor Statistical Indices in the Evaluation of Psychological Measures · Journal of Personality Assessment · 2015 · 816 citations
- Factorial Invariance Within Longitudinal Structural Equation Models: Measuring the Same Construct Across Time · Child Development Perspectives · 2010 · 621 citations
- A Comparison of Psychometric Properties and Normality in 4-, 5-, 6-, and 11-Point Likert Scales · Journal of Social Service Research · 2011 · 548 citations
- Estimating and interpreting latent variable interactions · International Journal of Behavioral Development · 2014 · 539 citations
- The NEO–PI–3: A More Readable Revised NEO Personality Inventory · Journal of Personality Assessment · 2005 · 479 citations
- Chi‐square for model fit in confirmatory factor analysis · Journal of Advanced Nursing · 2020 · 459 citations
- The Problem with Having Two Watches: Assessment of Fit When RMSEA and CFI Disagree · Multivariate Behavioral Research · 2016 · 415 citations
- Predictive model assessment and selection in composite-based modeling using PLS-SEM: extensions and guidelines for using CVPAT · European Journal of Marketing · 2022 · 406 citations
Most recent work
- Decoding Item Fit Statistics in Generalized Partial Credit and Graded Response Models · Measurement and Evaluation in Counseling and Development · 2026
- Technology-based versus paper-pencil: sources of mode effects in large-scale assessment · International Journal of Mathematical Education in Science and Technology · 2026
- Using parametric logistic item response theory model trees to detect item parameter heterogeneity in a high-stakes reading comprehension test · International Journal of Testing · 2026
- Review of automated parallel test form assembly · Behaviormetrika · 2026
- The Group Cohesion Scale-Revised: Reliability and Validity · The Journal of Psychodrama Sociometry and Group Psychotherapy · 2026
- A Step-by-Step Guide for Assessing DIF Using the Rasch Model · Measurement and Evaluation in Counseling and Development · 2026
- A Step-by-Step Tutorial to Conducting and Interpreting Measurement Invariance Testing for Counselors and Assessment Professionals · Measurement and Evaluation in Counseling and Development · 2026
- Consequential Validity-Centered Measure Development: An Illustration Using the ESSY Whole Child Screener · Journal of Psychoeducational Assessment · 2026
- Inconsistent Responding on a Mixed-Worded Scale in the PISA 2022 Questionnaire: Prevalence and Predictors Across Seven Countries · Journal of Psychoeducational Assessment · 2026
- Mapping First Grader’s Numerical Development at Scale: Leveraging Cognitive Models in a Large-Scale Educational Assessment · Educational Assessment · 2026
Find a gap in your own Psychometric Methodologies and Testing sub-topic
This page shows what the Psychometric Methodologies and Testing literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →