Education · 467 papers

Validation gaps in Education

617 open validation research questions in Educationgaps in reproducing, validating, or independently confirming findings — extracted from 467 papers in our local library. Below are representative open questions, each linked to the paper that raised it.

Representative open questions

Showing 30 of 617 — one per source paper, highest-quality first.

  • Power-Optimized AI-Enhanced Telepresence Robots: A Validated Multi-Modal Framework for Sustainable Remote Learning in Higher Education (2026) · doi

    The telepresence robot system was tested exclusively at 50 Mbps Wi-Fi connectivity; performance degradation at lower bandwidth speeds (e.g., 5 Mbps typical in rural classroom locations) remains uninvestigated, potentially overestimating NLP accuracy and latency metrics in resource-constrained educational environments.

  • El desarrollo de la comprensión del aspecto matemático del logaritmo mediante la intuición y la axiomatización en los estudiantes de Licenciatura en Educación Secundaria con especialidad en Matemáticas al término de la aplicación de su programa educativo. (2026) · doi

    The authors acknowledge that discrepancies exist between hand-calculated logarithmic results (using characteristic-mantissa separation) and calculator results, noting differences on the order of thousandths, but provide no systematic analysis of error accumulation patterns or guidelines for when hand-calculated methods are sufficiently accurate for pedagogical purposes.

  • e-DigCompEdu: Digital Competences for Online Higher Education Framework (2026) · doi

    The e-DigCompEdu framework aligns learners' digital competences with DigComp but lacks empirical validation of whether the five pedagogical competence areas (information/media literacy, digital communication/collaboration, digital content creation, responsible use, digital problem solving) effectively transfer to online higher education contexts with different disciplinary demands.

  • A systematic review of generative artificial intelligence in education: Pedagogical impacts, ethical risks, and future directions (2026) · doi

    The paper identifies fairness, prejudice, and transparency concerns in automated assessment systems using generative AI for essay grading and rubric creation, but does not specify empirical validation methods to measure bias across demographic student groups or establish benchmarks for acceptable fairness thresholds in AI-generated assessment feedback.

  • Ethical challenges of artificial intelligence in education: A systematic literature review on bias, privacy, and academic integrity (2026) · doi

    Multimodal learning analytics systems combine clickstream analysis, text analysis, speech recognition, gaze tracking, and biometric data to predict student dropout risk and performance, but the paper does not specify which combinations of these modalities are most effective for different student populations or learning contexts, nor what fairness-based machine learning validation frameworks should be applied to prevent bias in these predictive systems.

  • Error Patterns and Predictors in Solving Algebraic Word Problems: Evidence from Secondary School Students (2026) · doi

    Teachers report using algebra tiles and balance-scale manipulatives to address equation equivalence and inverse operations, but the paper does not quantify the effectiveness of these visual aids through experimental comparison or measure their impact on reducing process skill errors (reported at 67.5-87.5% frequency) across different student populations.

  • ChatGPT in computer programming education: A review of current literature and applications (2026) · doi

    Existing studies on ChatGPT in programming education employed predominantly small samples in qualitative and experimental designs, limiting generalisability of findings about learning outcomes and AI integration effectiveness. Large-scale, multi-institutional empirical studies are needed to validate reported benefits and challenges across diverse student populations.

  • Coding Science: The Role of Block-Based Programming Activities in Enhancing Computational Thinking and Science Achievement (2026) · doi

    Future research should develop integrated assessment tools specifically designed to evaluate science achievement through the lens of coding and iterative problem-solving, replacing traditional assessments like the FMKt that may not capture dynamic systems understanding fostered by block-based programming activities.

  • When teachers learn by observation: Vicarious reinforcement and student response systems in teacher education (2026) · doi

    While the study reports that QR code functionality malfunctioned during presentations, requiring manual data entry and disrupting interactive flow, there is no systematic analysis of how technical failures in student response systems affect the timing and effectiveness of vicarious reinforcement delivery or long-term adoption intentions among teacher educators.

  • A(n) (re)awakening: preservice teachers connecting their culture to science teaching along the US–México border (2026) · doi

    The study's small purposive sample of six Latina/o preservice teachers from two border universities limits transferability of findings to other cultural groups and non-border contexts; research with larger, more diverse samples across multiple geographic regions and ethnic backgrounds is needed to test the Community Cultural Wealth framework in science teacher preparation.

  • When faces trigger feelings: multi-stakeholder experiences, use, and expectations of facial-emotion learning technologies and materials for autistic people (2026) · doi

    The paper identifies intermediate or mixed emotional states (e.g., between happiness and sadness) as challenging for some autistic users to recognize in facial-emotion materials. Current tools focus on extreme expressions; there is a need to systematically test recognition and learning outcomes with facial-emotion stimuli displaying subtle blends and intermediate emotional expressions.

  • Segmental and suprasegmental pronunciation training via AI: An exploration of university students’ perceptions and attitudes (2026) · doi

    While the study explores student attitudes toward AI pronunciation tools, it lacks investigation of long-term retention effects and whether initial positive perceptions of AI-based segmental and suprasegmental pronunciation training sustain after 6-12 months of use.

  • The current landscape of teachers artificial intelligence acceptance: Relationships with TPACK and technostress (2026) · doi

    The relationship between sustainable professional learning ecosystems (collaborative communities, reverse mentoring) and both technopedagogical readiness and technostress reduction in AI adoption contexts has not been empirically validated through longitudinal or experimental designs.

  • Effect of Van Hiele group guided-discovery instructional approach on student engagement in learning plane geometry (2026) · doi

    The control group using traditional methods showed only minor improvements in cognitive and behavioral engagement components, but minimal data is provided on what specific traditional instructional practices were used. Future research should systematically compare VHGGDIA against multiple traditional geometry teaching approaches (direct instruction, worked examples, discovery without Van Hiele scaffolding) to isolate the contribution of the Van Hiele framework.

  • On the development and potential influential aspects of prospective science teachers’ competencies for the use of explainer videos for learning in physics education (2026) · doi

    The illusion of understanding phenomenon identified in first-year physics students using explainer videos has not been systematically investigated in prospective teachers; direct empirical testing of whether prospective teachers experience similar illusions of understanding requires comparative study designs with comprehension assessments.

  • Pluriliteracies for global citizenship: the 4Rs framework for deeper learning in the modern language(s) classroom (2026) · doi

    The paper proposes the 4Rs framework (Reading, Repositioning, Reflecting, Responding) as the language-as-discipline equivalent of DOEA, but does not empirically validate how these four practices operate distinctively in modern language classrooms compared to natural sciences or history classrooms where DOEA has been systematized. Comparative empirical studies are needed to specify the characteristic genres and activity domains through which language learners construct disciplinary knowledge.

  • Adequacy and Utilization of Physical Resources: Implications for Effective Teaching and Learning (2026) · doi

    While the study measured resource adequacy and utilization in public secondary schools using the AURTLOQ questionnaire and WAEC checklist, it did not assess whether these instruments are equally valid for detecting resource gaps in extremely remote or newly-constructed schools. Validation research is needed to determine instrument sensitivity across different school infrastructure age categories and geographic remoteness levels.

  • Digital Disruption of Academic Integrity (2026) · doi

    The study was limited to postgraduate students at public universities in Kenya; the generalizability of the IT mediation effects on the plagiarism-awareness-integrity relationship to undergraduate populations, private institutions, or different geographic contexts with varying technology access and policy frameworks requires cross-cultural and cross-institutional validation.

  • The Role of Artificial Intelligence in the Career Expectations of Ukrainian Students: Implications for Higher Education (2026) · doi

    The survey instrument was administered in English without formal translation or validation for the Ukrainian subsample, introducing language-related measurement error that may affect item comprehension and response comparability. A validated Ukrainian-language version of the measurement instrument is needed to ensure accurate assessment of ChatGPT Opportunity Index and Threat Index scores among Ukrainian student populations.

  • The Effects of Handwriting and Typing on Chinese Character Learning in the Digital Age: Evidence from Arab Adolescent Learners (2026) · doi

    While the paper found that Arab adolescents' semantic knowledge relied more on classroom instruction than sensorimotor channels (contradicting prior studies with other learner populations), no investigation examined whether this pattern holds for other non-Asian L1 groups or is specific to Arabic script users. Cross-linguistic validation with learners from different writing system backgrounds is needed.

  • No one-size-fits-all: a study of prompt techniques and large language models to enhance AI’s mathematics educational quality (2026) · doi

    The paper identifies that RAG achieved 99% explicit citation rates but cannot determine whether the cited Pólya heuristics were actually accessed versus hallucinated by the LLM. Future work should develop verification mechanisms to distinguish between genuine knowledge retrieval and citation hallucination in RAG-based mathematics educational systems.

  • Cluster Pattern Analysis of Students Stress using Machine Learning Algorithms with Feature Engineering (2026) · doi

    The three identified student stress clusters (high-risk with undesirable behaviors, moderate-risk, low-risk from stable families) are characterized qualitatively, but targeted intervention effectiveness has not been validated empirically. Future work should include longitudinal intervention studies measuring whether the recommended therapies and socialization strategies for each cluster actually reduce stress levels and improve retention in the identified student populations.

  • Multimodal AI in education: an avatar-based intelligent learning system for the Kazakh language (2026) · doi

    Gesture-speech synchronization is described as aligned with cultural communication patterns for Kazakh, but the paper does not provide empirical validation of whether the selected gestures and facial expressions are culturally appropriate or pedagogically effective for Kazakh learners compared to other language groups.

  • The use of large language models to solve mathematical modeling problems: preservice mathematics teachers’ use practices, perceived affordances and challenges, and trustworthiness judgments of AI-generated outputs (2026) · doi

    While the research documents that PSTs prefer ChatGPT over Thetawise due to familiarity and ease of use, there is no empirical comparison of how different LLM tool affordances (document handling, visual support, prompt interface design) systematically influence modeling phase engagement and output accuracy across different tool types.

  • Pre-service Teachers' Self-Directed Learning in a Blended Learning Environment: A Study on Scale Development and Affecting Factors (2026) · doi

    The SDL scale was validated exclusively at HNUE, one of seven key Vietnamese pedagogical universities; cross-cultural validation is needed across different countries and non-key institutional contexts to determine whether environmental and cultural differences significantly influence SDL levels in blended learning environments for pre-service teachers.

  • Development of User Story and Design Thinking Integration Teaching Model for Software Engineering Education (2026) · doi

    The model was evaluated over a limited number of learning sessions within a constrained timeframe, which may have artificially suppressed the perceived effectiveness of the US-DT teaching model and prevented students from mastering complex Design Thinking concepts. Future research must systematically test varying session durations and spacing strategies to determine optimal learning duration for achieving sustained improvements in all TAM factors including PEOU.

  • Does “Maniplo-Spatial Arty” Really Work? A Study of the Effectiveness of “Maniplo-Spatial Arty” as a Teaching and Learning in Science Pedagogy (2026) · doi

    Of 68 remembered maniplo-spatial tasks, 27 (40%) were recalled primarily for distinctive sensory factors (visual, acoustic, olfactory) and 18 (26%) for unusual presentation contexts, yet no systematic comparison exists between learning outcomes for maniplo-spatial tasks with high sensory/contextual salience versus standard laboratory-based maniplo-spatial tasks matched for conceptual content.

  • Quality Evaluation of Indonesian Student-generated User Stories: Insight from Human and ChatGPT Evaluation (2026) · doi

    ChatGPT achieved moderate-to-high agreement with human evaluators for structurally defined criteria but showed consistent lower agreement (55-65%) for context-dependent criteria like Unambiguous; future work must evaluate other AI models (beyond ChatGPT) and different prompting strategies to identify whether improved agreement for semantic user story quality assessment is achievable through model selection or instruction refinement.

  • Engagement in Exploratory Action Research to Improve Teaching Practice in EFL Writing (2026) · doi

    The study documents that personally relevant writing topics improved students' attitudes toward EFL writing, but lacks empirical quantification of how topic relevance (measured by student interest ratings or engagement metrics) correlates with actual writing performance gains. Research should employ mixed-methods designs measuring both affective shifts and linguistic/compositional outcomes in relation to topic relevance in EFL writing classrooms.

  • Washington University in St. Louis School of Medicine (2010) · doi

    While the paper mentions a sophisticated student-driven course evaluation system and collection of AAMC Matriculating and Graduation Questionnaires, USMLE exam performance, and residency director feedback, it does not describe specific psychometric validation studies or reliability/validity testing of these programmatic outcome measures across different student cohorts.

Working on one of these gaps? Publish with us.

Science AI Journal reviews manuscripts in under 15 minutes with 8 specialised AI reviewers calibrated on 69,000+ real peer reviews. Open access, CC BY 4.0.

Other gap types in Education

Command palette

Jump anywhere, run any action.