Open research questions in Hand Gesture Recognition Systems
34 unresolved questions extracted from the limitations and future-work sections of 251 Hand Gesture Recognition Systems papers in our library. Each links back to the study that raised it.
What the literature leaves open
References Advancements in machine learning (ML) and deep learning (DL) have garnered significant attention from researchers in the domain of sign language recognition (SLR). This study presents an extensive review of computational meth- odologies applied to SLR from 2003 to 2023. Drawing upon 988 publications sourced from the SCOPUS data- base through relevant keyword searches, the review offers a detailed examination of critical SLR processes, includ- ing image acquisition, segmentation, feature extraction, and classification. Techniques such as edge detection and skin color segmentation have demonstrated robust perfor- mance in segmentation, while hybrid approaches to feature extraction have shown promise in generating more reliable recognition features. The analysis highlights the necessity of high-performance computing systems to manage the intensive data processing demands associated with deep learning approaches. While significant progress has been made in recognizing words, alphabets, and numerals, the need for more research focused on sentence-level recogni- tion in sign language remains evident. Furthermore, since many intelligent SLR systems are still in experimental or prototype stages, it is essential to transition these models into practical applications for real-world use. Future directions for SLR research should prioritize enhancing accuracy and robustness through advanced ML and DL methods, particularly leveraging ensemble learning and transformer-based architectures. Efforts should also focus on improving image acquisition techniques, enhanc- ing the quality of input images, and exploring innovative image preprocessing approaches. Employing DL-based feature extraction methods can increase the discrimina- tive capabilities of SLR systems. Additionally, integrating multimodal information, such as body posture and facial expressions, could provide a more comprehensive under- standing of sign language and further enhance SLR system performance. Author Contributions Umang Rastogi (UR): UR has prepared the com- plete manuscript and done analysis of the published article. Rajen- dra Prasad Mahapatra (RPM): RPM has shared the suggestions to improve the manuscript. Sushil Kumar (SK): SK has planned for the preparation of manuscript and reviewed the manuscript thoroughly. Data Availability No datasets were generated or analysed during the current study.
Future research will focus on expanding the dataset beyond 200 words, incorporating multi-modal facial expressions and hand trajectories, and exploring hybrid Transformer–YOLO architectures to enhance contextual understanding and robustness in real-world scenarios.
Although Voice Activity Projection (VAP) has been successfully used to model future voice activity in spoken interaction, it remains unclear whether the framework transfers to sign language interaction.
Toward Signing Activity Projection in Sign Language Interaction · 2026Future research should focus on implementing real-time systems with actual radar hardware, incorporating temporal dynamics, and expanding the gesture vocabulary to enable more complex UAV operations.
Radar-based gesture recognition simulation for unmanned aerial vehicles command interpretation · 2026 · DOIFastSpeech2 earlier autoregressive Text-to-Speech (TTS) models by generating mel- spectrograms in parallel rather than sequentially. This significantly reduces inference time while maintaining high synthesis quality. The model incorporates variance predictors that explicitly model prosodic features such as pitch, duration, and energy, enabling more natural speech generation (Ikeda and Markov, 2024). During training, the FastSpeech2 model learns the mapping between textual input and corresponding speech signals using paired audio-text data. The text output produced by the sign language translation module is converted into phoneme or character representations and processed by the FastSpeech2 encoder. The model then generates a mel-spectrogram, which is subsequently converted into a waveform using a neural vocoder. By integrating FastSpeech2, the system can produce intelligible and natural-sounding Kazakh speech corresponding to the translated sign language sentences (Diatlova and Shutov, 2023). For automatic speech recognition (ASR), a Transformer-based encoder-decoder architecture was used, consisting of an 18-layer encoder and a 6-layer decoder with 8 attention heads and a hidden dimension of 512. Input features were 80-dimensional log Mel filterbanks augmented using SpecAugment, including time warping, frequency masking, and time masking. The model was trained using a hybrid CTC-attention loss with a CTC weight of 0.3, optimized with Adam and a warmup learning rate scheduler with 25,000 warmup steps. During inference, beam search with joint CTC-attention scoring was used for decoding. The ASR model processes audio input by extracting acoustic features from the speech signal and mapping them Frontiers in Artificial Intelligence 04 frontiersin.org Zhassuzak et al. 10.3389/frai.2026.1835419 FIGURE 3 The working principle of ASR algorithm. representations to corresponding textual (Figure 3). Speech signals are first transformed into spectral features that capture the temporal and frequency characteristics of the audio. These features are then processed by a neural recognition model that predicts the most probable sequence of text tokens. The recognized text is then used as input for the sign language translation pipeline, allowing spoken messages to be conveyed to hearing-impaired users.
Bidirectional Kazakh Sign Language prosody-aware translation using computer vision and speech recognition techniques · 2026 · DOI3 Future directions Future work will focus on expanding the dataset with addi- tional gestures and users, extending recognition to phrases beyond isolated words, enhancing the model to handle real- world video inputs, exploring alternative architectures and learning approaches, and validating performance on other sign language datasets, including Arabic Sign Language, to improve generalization and practical applicability.
An advanced model for gesture recognition in Indian sign languages classification-based accuracy improvement demonstrating robustness to rotation and scaling using deep learning approach · 2026 · DOIConclusion This research presented a real time sign language recognition system designed to assist communication between deaf and mute peoples. The proposed system uses a mobile based platform developed with flutter and integrated deep module for learning [26-30]. The backend through Node.JS and frontend through MongoDB for recognition of ASL & ISL sign language images and convert them to voice and text output for quiz learning platform. Shown in Figure 3.
• The system can be extended to support language recognition continuous sign IRJAEM 995 International Research Journal on Advanced Engineering and Management https://goldncloudpublications.com https://doi.org/10.47392/IRJAEM.2026.0151 e ISSN: 2584-2854 Volume: 04 Issue: 04 April 2026 Page No: 993 - 998 • sentences of ASL & ISL. Integration of voice and text output can allow to understand properly for better communication. • Developing energy-efficient architectures for low-cost devices AI • Model can be trained with larger and more diverse datasets to be converted for better voice communication. Integrating quiz-based learning platform of ASL & ISL sign images for better understanding, • • The system can be integrated with smart wearable devices.
The wireless communication integration (Bluetooth/Wi-Fi) for remote data transmission is proposed without addressing latency requirements for real-time sign language conversion or interference management in multi-user scenarios. Synchronization between flex sensor acquisition, wireless transmission, and remote text-to-speech generation needs protocol specification.
Smart Gloves for Sign Language to Text Conversion Translation Using Arduino Mega With 8 Flex Sensors · 2026 · DOIMiniaturization and wearability improvements are proposed, but the current system's form factor constraints are not analyzed. Research should quantify how adding accelerometers and gyroscopes affects glove bulk, user comfort during extended wear, and signal noise from motion artifacts, with validation on actual users over specified wear periods.
Smart Gloves for Sign Language to Text Conversion Translation Using Arduino Mega With 8 Flex Sensors · 2026 · DOIPower consumption is listed as 'moderate' without quantitative specification (mAh, hours of continuous operation, battery requirements). A detailed power budget analysis comparing flex sensor polling frequency, Arduino Mega processing cycles, DF Player audio amplification, and LCD display draw is needed to enable portable battery-powered deployment.
Smart Gloves for Sign Language to Text Conversion Translation Using Arduino Mega With 8 Flex Sensors · 2026 · DOIProposed expansion to support multiple languages is mentioned without specifying how gesture-to-language mapping would be implemented for sign languages with different phonetic structures. Research is needed on whether a single hardware configuration can distinguish between sign language variants (e.g., ASL vs. BSL) or if separate calibration profiles are required for each language.
Smart Gloves for Sign Language to Text Conversion Translation Using Arduino Mega With 8 Flex Sensors · 2026 · DOIThe DF Player Mini module currently relies on pre-recorded audio files, limiting linguistic flexibility and natural language output. Implementation of embedded text-to-speech (TTS) systems on the Arduino Mega platform needs investigation, including memory constraints, processing latency, and support for multiple languages simultaneously on resource-limited microcontrollers.
Smart Gloves for Sign Language to Text Conversion Translation Using Arduino Mega With 8 Flex Sensors · 2026 · DOIThe integration of machine learning algorithms for dynamic and adaptive gesture recognition is proposed as future work, but no specific algorithm (e.g., k-NN, SVM, neural networks) or training dataset characteristics are defined. Research should specify which machine learning models are suitable for Arduino Mega's computational constraints and how sensor data should be preprocessed for model training.
Smart Gloves for Sign Language to Text Conversion Translation Using Arduino Mega With 8 Flex Sensors · 2026 · DOIThe paper demonstrates gesture recognition using threshold-based algorithms with 8 flex sensors on a single gesture set, but does not specify the dataset size, number of unique gestures tested, or validation across multiple subjects. A quantitative evaluation with defined gesture vocabularies (e.g., ASL alphabet, common phrases) and cross-subject validation is needed to establish generalizability.
Smart Gloves for Sign Language to Text Conversion Translation Using Arduino Mega With 8 Flex Sensors · 2026 · DOIThe smart glove system currently addresses sensor placement inconsistencies through manual calibration, but no systematic methodology is provided for standardizing flex sensor positioning across different hand sizes or morphologies. A calibration protocol that accounts for anatomical variation in finger length and joint positioning should be developed to ensure consistent gesture recognition accuracy.
Smart Gloves for Sign Language to Text Conversion Translation Using Arduino Mega With 8 Flex Sensors · 2026 · DOIThe system implements eye landmark detection for zooming and file operations using MediaPipe, but no threshold values, sensitivity parameters, or eye movement classification criteria are documented, making reproducibility and tuning for different users unclear.
The modular architecture supports adding new gestures (e.g., swipe for navigation) without specifying how gesture confusion matrices or inter-gesture discrimination metrics would be measured when the gesture vocabulary expands beyond the currently implemented set.
User-centric testing and iteration is mentioned as part of the development strategy, but the paper does not specify the sample size, demographic diversity, gesture learning curve, or adaptation time required for diverse users to achieve proficiency with the gesture control mappings.
The system is designed for accessibility with basic webcams, but performance variation across different webcam resolutions, frame rates, and quality levels (built-in vs. external cameras) has not been tested or characterized, creating uncertainty about minimum hardware requirements.
Real-time performance optimization for frame processing pipelines is discussed generally, but no latency benchmarks, frame-per-second measurements, or comparison of lightweight algorithms versus standard approaches are provided for the gesture detection pipeline on the specified hardware (Intel i5, 8GB RAM).
The human activity recognition component (detecting activities like 'Talking' or 'Laughing') is mentioned as future extensibility but lacks specification of which pose estimation landmarks or facial feature detection methods would be required to distinguish these activities from MediaPipe's pose and face mesh outputs.
The paper implements debounce logic and thresholding mechanisms to minimize false triggers in gesture detection, but provides no empirical evaluation of false positive/negative rates across different gesture types or user populations. Quantitative performance metrics for error handling reliability are absent.
This project manages to introduce a convenient, dependable and economical smart home automation system based on gesture control, combined with a timer and mobile surveillance. It enables the system to benefit from real-time hand gesture recognition which is accurate and responsive by using MediaPipe and OpenCV. The addition of time-based automation and visualization in the form of a mobile app goes a long way towards enhancing convenience, safety, and accessibility to the users. The system has a high potential of practical implementation in smart living. The further progress of the product can be aimed at the application of machine learning-enabled adaptive gesture recognition, voice-controlled assistants, expansion of control over multi-appliances, and the introduction of cloud-based intelligent automation. the modular [9] In addition, architecture that provides the ease of customization according to various home settings without making significant changes in hardware or software. Raspberry Pi as the central processor unit is also used to ensure that the power consumption is low, and at the same time, Raspberry Pi provides enough performance to operate in real-time. The mobile monitoring option increases the reliability of the systems as it would give regular updates on the status even when limitations are experienced. The project will provide a strong base of scalable and user-friendly smart home solutions with a focus on interaction, energy intuitive efficiency, and continuous connectivity. Future upgrades could be a safe cloud connection to allow for remote access, customized learning algorithms to enhance accuracy, and a smooth integration with the currently existing IoT platforms.
Testing involved only 20 hand movements from a balanced set of male and female subjects with 60 images per gesture; cross-subject validation and cross-dataset evaluation (e.g., testing on external datasets) are absent, limiting generalization assessment for the hand gesture recognition system.
Most-cited papers in Hand Gesture Recognition Systems
- mm-Pose: Real-Time Human Skeletal Posture Estimation Using mmWave Radars and CNNs · IEEE Sensors Journal · 2020 · 367 citations
- Capturing complex hand movements and object interactions using machine learning-powered stretchable smart textile gloves · Nature Machine Intelligence · 2024 · 124 citations
- A deep learning approach for evaluating the efficacy and accuracy of PoseNet for posture detection · International Journal of Systems Assurance Engineering and Management · 2024 · 85 citations
- Multi-task learning for hand heat trace time estimation and identity recognition · Expert Systems with Applications · 2024 · 66 citations
- Real-Time Arabic Sign Language Recognition Using a Hybrid Deep Learning Model · Sensors · 2024 · 61 citations
- A survey on hand gesture recognition based on surface electromyography: Fundamentals, methods, applications, challenges and future trends · Applied Soft Computing · 2024 · 59 citations
- A Survey on Yogic Posture Recognition · IEEE Access · 2023 · 48 citations
- Sign language recognition using artificial intelligence · Education and Information Technologies · 2022 · 36 citations
- Detection of hand gestures with human computer recognition by using support vector machine · Periodicals of Engineering and Natural Sciences (PEN) · 2022 · 27 citations
- Smart Assist System Module for Paralysed Patient Using IoT Application · EAI Endorsed Transactions on Internet of Things · 2024 · 25 citations
Most recent work
- AI-Powered Smart ISL Translator with Voice, Text & Gesture Recognition · IJIREEICE · 2026
- Machine Learning-Based Vision-to-Speech System for Assistive Application · International Journal for Research in Applied Science and Engineering Technology · 2026
- Vision and Voice Controlled Bionic Robotic Arm for Assistive and Research Applications · International Journal of Creative and Open Research in Engineering and Management · 2026
- Real-Time Hand Control for Interactive Presentation · International Journal for Research in Applied Science and Engineering Technology · 2026
- Human Activity Recognition · International Journal for Research in Applied Science and Engineering Technology · 2026
- MOVE, PLAY, ENGAGE - A REAL-TIME MOTION CONTROLLED GAMING PLATFORM · International Scientific Journal of Engineering and Management · 2026
- Gesture-Controlled Home Automation with Mobile Interface for Disabled and Elderly · International Research Journal on Advanced Engineering Hub (IRJAEH) · 2026
- Smart Gloves for Sign Language to Text Conversion Translation Using Arduino Mega With 8 Flex Sensors · Research Digest on Engineering Management and Social Innovations · 2026
- Ai-Based Hospital Assistance System Using Indian Sign Language Translation · Zenodo (CERN European Organization for Nuclear Research) · 2026
- Air Writing Recognition · International Scientific Journal of Engineering and Management · 2026
Find a gap in your own Hand Gesture Recognition Systems sub-topic
This page shows what the Hand Gesture Recognition Systems literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →