Computer Science · Research topic

Open research questions in Human Pose and Action Recognition

76 unresolved questions extracted from the limitations and future-work sections of 278 Human Pose and Action Recognition papers in our library. Each links back to the study that raised it.

What the literature leaves open

  • Large rotations, rapid motion, and frequent occlusion in sports impacts. Limited accuracy of standard head pose benchmarks for head acceleration event conditions. Difficulty in estimating head pose during large, multi-axis rotations.

    Deep learning approaches for head pose estimation in sports impacts · 2026 · DOI
  • The study only tested the models on a dataset of controlled football headers. The dataset was limited to 10 participants. The study did not evaluate the models on other types of head acceleration events.

    Deep learning approaches for head pose estimation in sports impacts · 2026 · DOI
  • The need for trustworthy AI systems to support caregivers in kindergartens. The lack of effective monitoring systems for ensuring child safety.

    Real-Time Monitoring of Kindergarten Safety Using YOLO-11-Based Detection of Children and Adults · 2026 · DOI
  • Large computational load of human pose estimation models. Slow inference speed of human pose estimation models. Limited hardware resources for deployment.

    A Lightweight Human Pose Estimation Algorithm Based on Improved YOLO11-Pose · 2026 · DOI
  • Challenging imaging conditions due to lens soiling and grid occlusions. Limited field of view and lack of overlap between cameras. High variability in behavioral classification performance.

    Automated behavioral segmentation and markerless pose tracking of mice during spaceflight · 2026 · DOI
  • The demand for multi-person, real-time pose estimation has surged across various domains. There is a need for a detailed benchmarking of state-of-the-art frameworks.

    Multi-person 2D human pose estimation a benchmark for real-time applications · 2026 · DOI
  • Further studies could be conducted to evaluate the system on a larger sample of participants. The system could be integrated with other technologies to provide real-time feedback to users.

    Deep learning-based physical exercise assessment of older adults using single-camera videos · 2026 · DOI
  • There is a lack of personalized supervision for care-home residents to engage in physical activity. Autonomous, technology-supported exercise platforms could fill this gap.

    Deep learning-based physical exercise assessment of older adults using single-camera videos · 2026 · DOI
  • Current methods cannot meet the needs for higher-precision identification of instrumental motions. Prior work on human pose estimation technology in music performance analysis has limitations.

    Design of University Instrumental Music Performance Movement Recognition and Accurate Feedback System Integrating OpenPose and LSTM · 2026 · DOI
  • The gap in rehabilitation training evaluation is the lack of effective models that can provide real-time feedback and guidance. Coarse feedback granularity and high labeling costs are common challenges in rehabilitation training evaluation.

    <b>Research on Rehabilitation Training Movement Recognition and Real-time Feedback Model Based on Computer Vision</b> · 2026 · DOI
  • Creating and validating accurate tracking models is time-consuming and labor intensive. Many research groups duplicate efforts on similar images.

    Automated behavioral tracking of zebrafish larvae with DeepLabCut and SLEAP: pre-trained networks and datasets of annotated poses · 2026 · DOI
  • A key open problem, however, is how to learn triplet representations that remain reliable across institutions, where surgical video varies in acquisition conditions, surgeon style, tool usage, and tissue handling, while existing triplet datasets do not support explicit evaluation of center-wise transfer.

    SPIRIT: Spatio-temporal Pairwise Relational Modeling of Instrument-Tissue Interactions for Surgical Action Triplet Recognition · 2026
  • Knowledge Distillation (KD) offers a promising yet underexplored path for compressing large action recognition models.

    Channel-wise Dynamic Knowledge Distillation via Adaptive Sample Generation for Action Recognition · 2026
  • However, existing methods are generally limited to a single device, which restricts the effective monitoring range.

    PressureMesh: 3D Human Mesh Estimation from Multi-Device Pressure Images · 2026
  • Future studies will focus on lightweight ViTs, multimodal data fusion, and self-supervised learning to facilitate the implementation of these approaches in various areas, including smart surveillance, healthcare monitoring, AR/VR interfaces, and human-robot collaboration in smart environments.

    ViT-HAR: Vision Transformer-Based Human Activity Recognition in Cluttered Environments · 2026 · DOI
  • Future studies will focus on lightweight ViTs, multimodal data fusion, and self-supervised learning to facilitate the implementation of these approaches in various areas, including smart surveillance, healthcare monitoring, AR/VR interfaces, and human-robot collaboration in smart environments. Although a lack of data is one of the problematic issues, the next generation of the ViT-HAR system will include robustness analysis in restricted data conditions.

    ViT-HAR: Vision Transformer-Based Human Activity Recognition in Cluttered Environments · 2026 · DOI
  • To address this limitation, future work will focus on expanding the dataset to include a broader spectrum of postures and gestures and on exploring more advanced learning strategies to improve robustness and generalization. Building upon our prior work on eye-based visual attention and fatigue analysis [31], facial geometry–based engagement detection[32], and EEG-based cognitive state modeling [33], future research will investigate multimodal engagement analytics by fusing pose-based cues with ocu- lar, facial, and EEG/BCI signals.

    Human pose estimation based engagement detection in E-learning: a hybrid approach · 2026 · DOI
  • acknowledges that contextualize these contributions and suggest directions for future work. The proposed biomechanical engine employs a 2D planar arm model, a reasonable simplification for frontal gesture analysis but inadequate for actions with significant depth variation or out-of-plane motion. Extension to full 3D biomechanics would require depth sensing (stereo cameras, LiDAR) or multi-view inputs, enabling modeling of complex 3D joint rotations and whole-body coordination. The proposed physics model captures essential dynamics through a 2-link arm with standard inverse dynamics, but more sophisticated biomechanical models could incorporate muscle activation patterns, multi-segment coordination, contact forces, or fatigue effects—potentially offering even richer physical signals for complex action recognition tasks. The computed kinetic features (torque and energy) are most informative for gestures that involve significant limb motion with clear physical signatures. For subtle actions like facial expressions or fine finger movements where forces are minimal, physics-based features may provide limited additional information beyond geometric pose features. The evaluation focused primarily on a single domain— police traffic gestures—demonstrating the feasibility and benefits of physics-informed learning in a controlled setting. While this establishes proof of concept, broader validation across diverse action recognition benchmarks (e.g., NTU RGB+D for full-body actions, Kinetics for everyday activities, domain-specific datasets in sports or sign language) would strengthen claims about generalization and reveal which action www.etasr.com Haimer et al.: Physics-Informed Deep Learning for Human Action Recognition: A Biomechanical … Engineering, Technology & Applied Science Research Vol. 16, No. 2, 2026, 33854-33865 33864 categories benefit most from physics integration. The physics computation adds approximately 12% inference overhead compared to pure pose-based classification; for ultra-low- latency physics computation (e.g., periodic rather than per-frame calculation) could reduce this cost while maintaining benefits. investigating applications, selective Several promising directions emerge for extending this work. Multi-modal physics integration could combine visual pose estimation with Inertial Measurement Units (IMUs), force plates, or Electromyography (EMG) sensors, providing ground- truth physical measurements for even stronger supervision and enabling applications in sports biomechanics or rehabilitation monitoring. Hierarchical physics modeling could extend from isolated limbs to full-body coordination, capturing higher-level principles, such as center of mass dynamics, momentum conservation, balance maintenance, and whole-body energy optimization—relevant for complex activities such as dancing, martial arts, or gymnastics.

    Physics-Informed Deep Learning for Human Action Recognition: A Biomechanical Approach · 2026 · DOI
  • The need for innovative solutions to master yoga poses, particularly for beginners lacking access to experienced instructors. The lack of lightweight and efficient models that can provide accurate pose estimation while conserving computational resources.

    Yoga Pose Estimation and Prediction · 2026 · DOI
  • The experimental evaluation was conducted primarily on the UCF101 dataset. The study does not fully capture the variability present in real-world environments. Real-time deployment on embedded hardware requires further validation.

    Efficient Human Action Recognition Using MobileNetV1 and EfficientNetB3 Based Hybrid Network · 2026 · DOI
  • Further validation through device-specific optimization and benchmarking is required. The study can be extended to other datasets and real-world environments.

    Efficient Human Action Recognition Using MobileNetV1 and EfficientNetB3 Based Hybrid Network · 2026 · DOI
  • Conventional vision-based posture identification algorithms encounter issues such as occlusions and background clutter. IMU-only techniques are inefficient in detecting misalignments such as nonuniform joint extension or shoulder tilt.

    Hybrid vision–IMU deep learning framework with graph convolutional networks and attention for personalized yoga posture identification · 2026 · DOI
  • In this study, a hybrid attention-driven deep learning architecture is used to combine vision-based skeleton key points with IMU sensor data to produce a personalised yoga posture recognition system. The proposed model effectively employs an attention-based fusion module to produce consistent and user-adaptable representations, with GCN for vision features and LSTM for IMU signals. According to experimental results on the Yoga-82 dataset and a bespoke IMU dataset, the hybrid model outperforms the current vision-only, IMU-only, and static fusion models on all major parameters, achieving 96.2% accuracy, which is 2-7% higher than previous models. Achieved a high F1-score (~95-96%), recall, and accuracy, indicating reliable multi-class yoga position memory. It has a low computational cost (37.5 GFLOPs) and fast inference (~27 ms per frame), making it suitable for real-time applications. The outstanding adaptation of the tailored learning module to human variances reduces the number of postures that appear to be identical but are misclassified. The results demonstrate how attention-based fusion and personalization in hybrid multi-modal systems may greatly improve the accuracy and robustness of fine-grained human posture recognition, which qualifies the technology for use in digital wellness, physiotherapy, and fitness monitoring systems. Using the current design, several methods may be used to further enhance the system, including: Real-Time Correction and Feedback: Customers can get real-time feedback by integrating posture correction assistance with biomechanical angle analysis. Develop lightweight transformers or quantized models that may be used on wearables and mobile devices to optimize edge deployment. 3D Pose and Depth Integration: Use multi-view cameras or depth sensors to get more precise spatial data to enhance occlusion control. Cross-Environment Generalization: To increase adaptability to various backgrounds and lighting conditions, extend ARTICLE IN PRESS ARTICLE IN PRESS ACCEPTED MANUSCRIPT training to many settings (studio, home, and outdoor). By encouraging users to gradually adopt more precise postures, reinforcement learning— which makes use of adaptive feedback loops—can assist users in improving their posture.

    Hybrid vision–IMU deep learning framework with graph convolutional networks and attention for personalized yoga posture identification · 2026 · DOI
  • Replacing all MLP components with KANs in HAR models often degrades accuracy and computation efficiency, highlighting an open challenge: how to combine KANs' precision with MLPs' noise robustness and efficiency.

    KAN-MLP-Mixer: A comprehensive investigation of the usage of Kolmogorov-Arnold Networks (KANs) for improving IMU-based Human Activity Recognition · 2026
  • Vision-based models are inherently sensitive to environmental changes such as lighting, occlusion, and camera viewpoint. Wearable systems require precise sensor placement and careful calibration, and may cause user discomfort. The lack of quantitative data compromises the consistency and precision of outcome evaluations.

    Spatio-temporal graph attention network for rehabilitation movement classification · 2026 · DOI

Most-cited papers in Human Pose and Action Recognition

Most recent work

Find a gap in your own Human Pose and Action Recognition sub-topic

This page shows what the Human Pose and Action Recognition literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.

Open the Research Gap Finder →

Related topics in Computer Science

76 open questions have been extracted from the limitations and future-work passages of 278 Human Pose and Action Recognition papers in our library. Each one below links back to the study that raised it, so you can read the original claim in context.

Tools for your next paper

Compare the categoryHonest roundups of the AI research tools, ours listed alongside the alternatives.

Command palette

Jump anywhere, run any action.