Psychology · Research topic

Open research questions in Emotion and Mood Recognition

99 unresolved questions extracted from the limitations and future-work sections of 519 Emotion and Mood Recognition papers in our library. Each links back to the study that raised it.

What the literature leaves open

  • 4 Challenges and Open Questions 6. 4 Challenges and Open Questions The THAI-SER corpus can play an important role in the research area of speech emotion recognition.

    Thai speech emotion recognition (THAI-SER) corpus · 2026 · DOI
  • Future research should consider developing a structured intervention protocol or response framework to ensure consistency in how instructors adjust their teaching strategies based on real-time emotion detection. Further, the involvement of only two instructors limits the generalizability of the findings, as teaching styles can vary widely.

    Development and Evaluation of a Real-Time Emotion Detection System to Enhance Student Interaction · 2026 · DOI
  • The current study has some limitations. First, although we were interested in comparing how variations in the dynamic characteristics of the stimuli might impact ER accuracy, the stimuli also varied in duration across tasks (see Supplemental Materials). These duration differences are due to the inherent properties of the emotional expressions of interest: full sentences (in the dynamic vocal condition) are naturally longer than the vocal bursts (in the static vocal condition), but shorter than the videos (in the dynamic facial condition). This is reflective of the types of nonverbal expressions people typically encounter in everyday life, offering a valuable opportunity to explore differences in individuals’ ability to interpret relevant social stimuli. Nonetheless, it is possible that longer stimuli may have required additional processing demands, heightening the relative difficulty of labelling their emotional intent. However, given that the longest stimulus type (dynamic faces) was also the best recognized, we estimate that differences in stimulus duration are not likely to fully explain differences in ER accuracy across tasks, nor associations with cognitive abilities. Second, while it was necessary to low-pass filter the dynamic vocal condition stimuli (to ensure no condition included linguistic content that could influence ER), this manipulation reduces the richness and clarity of acoustic cues and increases perceptual ambiguity. Consequently, the absence of semantic content, combined with the degraded acoustic signal, likely contributed to participants’ lower emotion recognition accuracy in this condition. Nonetheless, prior work has demonstrated that low-pass filtering retains the prosodic cues necessary for accurate emotion identification (Knoll et al., 2009; Snel & Cullen, 2013). Additionally, low accuracy rates for prosody compared to facial stimuli are common in the ER literature (reviews in Morningstar et al., 2018; Zupan & Eskritt, 2024). Thus, although absolute accuracy rates were likely suppressed for this condition, we do not think it likely that our findings of comparatively lower accuracy for the dynamic voice ER task were driven by low-pass filtering alone. To balance ecological validity and control for semantic content, one approach could be to incorporate semantic content into both the dynamic vocal and dynamic facial conditions, ensuring more equitable comparisons. However, integrating Journal of Nonverbal Behavior1 3 auditory components with visual stimuli in the dynamic facial condition would complicate the design by introducing multimodal elements. This fundamental distinction would set the dynamic facial condition apart from the other three unimodal conditions in the experiment.

    Emotion Recognition Accuracy Across Nonverbal Modalities: Associations with Working Memory · 2026 · DOI
  • Although the inclusion of EMODB alongside TESS provided a degree of acoustic and demographic variability, the generalisability of the findings remains limited.

    Deep Learning-Based Speech Emotion Recognition for IoT Edge Devices: A Comparative Study · 2026 · DOI
  • While recent advances in deep learning have significantly improved SER performance in Indo-European languages, Arabic SER remains underexplored and challenging due to dialectal diversity, limited annotated datasets, and the difficulty of modeling both local spectral cues and long-range temporal dependencies.

    Towards Robust Arabic Speech Emotion Recognition with Deep Learning · 2026
  • This research proposes a multiple-component-based framework for the analysis of student engagement. Facial emotion recognition, visual attention estimation, and reasoning about time were all incorporated into the unified framework. Unlike in the conventional black-box and single-modality-based methods, the components of the analysis of student engagement were transparent and cognitively informed. Facial emotions are first encoded using a dual-stream model consisting of a transformer-graph architecture. The global facial appearance semantics are encoded using a vision transformer based on the BEiT model. Simultaneously, the relational dependencies between the dense facial landmarks are encoded using a Graph Attention Network (GAT). The encoded representations are then fused to obtain robust emotion predictions for an individual’s frames. Subsequently, these predictions are mapped to the relevant affective states for engagement. Visual attention cues arising from gaze direction and eye openness are used to refine per-frame engagement predictions using a hierarchical fusion strategy. Temporal smoothing is then applied to make the engagement predictions more temporally consistent. This reduces sensitivity to microexpressions, gaze changes, and noise in landmark detection. The reasoning over multiple timescales allows the system to correctly identify patterns of engagement that are not subject to short-term visual artifacts. The effectiveness of the proposed framework can be comprehensively justified by the experimental analysis. It can be clearly evaluated that the ARTICLE IN PRESS ARTICLE IN PRESS ACCEPTED MANUSCRIPT implies balanced, strong performance in facial emotion framework recognition across all emotion classes. It can also be confirmed that the model is focusing on anatomically and psychologically important areas of the faces by means of the qualitative explainability analysis of the model, as it can solve the task considering the means of transformers and graph reasoning. In conclusion, the proposed system provides an unobtrusive, transparent, and practical solution for analyzing real-time student engagement in an elearning environment using regular webcams. Future extensions and further improvements to the system would focus on quantitative verification of results using available datasets on student engagement, learning decision thresholds, and the use of other modalities to understand engagement in learning scenarios. 'Declarations'. Methods: We confirm that our study did not involve any experiments on humans, human tissue samples, or animals. Therefore, ethical approval and informed consent were not applicable to this research.

    A hierarchical transformer–graph framework for explainable student engagement estimation in E-learning videos · 2026 · DOI
  • representation of rare affective states. In addition, systematic comparisons of embedding architectures, including multilingual and monolingual transformers, distilled and full-scale models, and non-transformer alternatives, should be conducted to assess semantic bias, robustness to paraphrasing, and transferability across domains. Progress toward comprehensive multimodal fusion involves synchronizing facial landmark data with speech prosody, linguistic transcripts, and physiological signals such as EKG, EEG, skin conductance, and eye tracking. It is important to investigate whether audio and visual features can predict missing physiological information in order to enable reliable emotion estimation in environments with limited sensor data. Finally, developing AFFECT as a realtime, user-friendly platform that is accessible through web APIs, desktop interfaces, or mobile applications, and that supports modular integration of multiple languages, theoretical models such as the circumplex and appraisal frameworks, and diverse application domains and marketing, will be crucial for practical deployment. Together, these improvements will enable more accurate, flexible, and context-sensitive emotion recognition in both research and applied settings.

    Automated emotion recognition via video-based semantic embeddings · 2026 · DOI
  • Although the present study was not designed for clinical diagnosis or deployment, the proposed multimodal framework may have potential relevance for future research on aect-aware monitoring systems. Emotional dysregulation is an important feature of several psychiatric conditions, including depression, anxiety disorders, and post-traumatic stress disorder, which explains the broader interest in objective physiological markers of aective state. The ability to objectively quantify emotional states using non-invasive physiological signals oers a pathway toward continuous and ecologically valid monitoring, complementing traditional symptom-based assessments that rely on episodic selfreporting (Wang and Wang, 2025). From an engineering perspective, the parallelizable structure of self-attention may be advantageous for processing longer the current study was physiological recordings. However, conducted in a controlled laboratory environment using healthy young adults and wired acquisition systems. Accordingly, the present framework should be interpreted as a proof-offor multimodal physiological emotion recognition concept rather than an immediately deployable healthcare solution. The parallelizable nature of self-attention enables eÿcient processing of long-duration recordings, facilitating integration into wearable or bedside monitoring systems. Moreover, the interpretability aorded by attention mechanisms enhances clinical trust by enabling practitioners to inspect temporally salient physiological patterns associated with emotional dysregulation. Such transparency is increasingly recognized as a prerequisite for the adoption of artificial intelligence tools in clinical practice (Metzner et al., 2024). Nevertheless, several limitations warrant critical consideration. In addition, the model was evaluated using subject-independent cross-validation within a single dataset, without an external validation cohort or cross-dataset testing. As a result, robustness across dierent recording environments, sensor configurations, and stimulus protocols remains to be established. Emotional expression and physiological responsiveness are known to vary across developmental stages and psychopathological conditions. Furthermore, although EEG and GSR provide complementary information, low-arousal emotional states such as Neutral and Sadness remained more diÿcult to distinguish, likely because these categories share relatively similar autonomic intensity while diering more subtly along valence-related dimensions. Further refinement is required to enhance sensitivity to lowarousal aective disturbances. Future research directions include expanding the multimodal repertoire to incorporate heart rate variability, respiration patterns, and facial electromyography, as well as validating the model in clinically diagnosed populations. Through such extensions, the proposed approach may contribute to the development of more generalizable and physiologically informed emotion recognition the comparative evaluation included canonical CNN and Bi-LSTM baselines but did not include more recent graph-based or hybrid Transformer architectures, which should be addressed in future work to strengthen state-of-the-art comparisons. systems. Furthermore, Compared with prior multimodal emotion recognition studies that rely primarily on feature concatenation or shallow fusion strategies, the present framework explicitly models temporal dependencies within and across modalities using attention mechanisms. In this sense, the study provides an additional perspective on central–peripheral interaction in multimodal physiological emotion recognition, while further validation is still needed across broader datasets and experimental contexts.

    Emotion recognition based on the temporal patterns of electroencephalogram signals and electrodermal response signals using the TRANSFORMER network · 2026 · DOI
  • Future research will focus on extending the framework to real-time FER systems and integrating multimodal emotion recognition using audio and physiological signals. Additionally, advanced transformer-based architectures can be incorporated enhance performance. Although the proposed to further model demonstrates strong performance, certain limitations remain. Performance may decrease under extreme occlusion investigate or transformer-based emotion recognition frameworks. low-resolution conditions. Future work will and multimodal architectures consistently achieved superior performance The overall results obtained from the experimental analysis clearly demonstrate the effectiveness and reliability of the proposed explainable gradient convolutional vector fuzzy pattern recognition (ExGrConVFuzPR) model for facial expression recognition. The model across multiple datasets, with notable accuracy improvements on JAFFE, CK, and AFLW, highlighting its adaptability to varied facial structures and lighting conditions. The visual segmentation results provided transparent insights into the feature extraction process, confirming that the explainable gradient layer accurately focuses on emotion-relevant regions such as the eyes, mouth, and eyebrows. Comparative evaluations further confirmed that ExGrConVFuzPR significantly outperformed existing methods including CNN, RF, fuzzy rules, DT, and GA, with measurable enhancements in accuracy, precision, recall, and error reduction. These findings validate that the integration of explainable gradient-based segmentation and fuzzy pattern reasoning leads to a robust and interpretable learning framework, establishing ExGrConVFuzPR as a promising approach for advancing explainable and highaccuracy facial expression recognition in real-world applications.

    Explainable gradient convolutional vector fuzzy pattern analysis based on ensemble model for facial expression recognition · 2026 · DOI
  • We proposed MoodSenseAI an adaptive, attention-guided, multimodal deep learn- ing system for depression detection that integrates and utilises semantic, acoustic, and behavioural cues across text, speech, and video modalities. This addresses a primary lim- itation of unimodal, static fusion approaches by enabling the proposed DeepMoodNet model to extract modality-specific features and dynamically fuse them via an attention mechanism. Validation of the Experimental results of our proposed framework sur- passes the various single- and bimodal baseline systems and achieves a Macro Accuracy of 94.3%. Ablation studies provide additional evidence of the necessity of each modality and the importance of comprehensive depression evaluations. This work addressed mul- tiple key limitations in the state of the art (modality imbalance, limited adaptive fusion, and lack of explainability). The modular and scalable design of MoodSenseAI enables its implementation in clinical telehealth platforms, social media mental health screen- ing, and monitoring systems for continuous screening across all environments, enabling automated detection of depression at the earliest. Our future work will focus on the limitations discussed in Sect. 5.1, including validation on larger, more diverse clinical datasets and improving the system’s robustness against missing or noisy modalities. We will build on this work for more sophisticated mental health assessments, such as multi- class classification of depression severity levels and longitudinal mental health monitor- ing. Also, achieving model-level transparency and modality contribution analysis with explainable AI methods and real-time adaptation in telehealth settings will be the pri- mary focus. Focusing on improving these will only help MoodSenseAI become a cred- ible and feasible tool for preventive mental health treatment and personalised attention. Future work will validate the framework under subject-aligned multimodal conditions using synchronized text–speech–video streams collected per individual in telehealth or clinical environments, and will examine robustness under missing-modality sce- narios. Future work will focus on handling noisy and missing modalities, incorporating subject-aligned multimodal datasets, and improving robustness for real-world clinical deployment.

    MoodSenseAI an attention-guided multi-source multimodal deep learning system for interpretable depression detection from text speech and video · 2026 · DOI
  • Future work will focus on extending the framework to larger and more diverse datasets, incorporating heterogeneous recording systems, and exploring bias- aware, domain-generalization, and temporal or channel-wise attention mechanisms to further enhance robustness and interpretability in EEG-based emotion recognition. Moreover, the analysis in this study is restricted to binary gender labels (male and female) as provided by the available datasets and does not account for nonbinary or gender-diverse identities, which should be considered in future EEG-based emotion recognition studies as more inclusive datasets become available.

    An Interpretable Multiscale Temporal Residual Framework for Gender-Specific EEG-Based Emotion Recognition · 2026 · DOI
  • Clinical Validation: The system has not been validated against a clinical reference standard or compared to gold-standard diagnostic interviews. Future work should include prospective validation studies comparing system outputs to clinician diagnoses using instruments such as the SCID or HDRS [1, 2, 3]. Model Generalization: The individual CNN and LSTM models were trained on specific emotion recognition datasets and may not generalize optimally to diverse populations with varying demographics, cultural backgrounds, or depressive presentations. Fairness and Bias: Facial recognition systems are known to exhibit performance disparities across demographic groups. Future work should analyze system performance across different populations and IJLRP26052191 Volume 7, Issue 5, May 2026 8 International Journal of Leading Research Publication (IJLRP) E-ISSN: 2582-8010 ● Website: www.ijlrp.com ● Email: [email protected] implement bias mitigation strategies. Temporal Dynamics: The current system treats each assessment as independent. Integrating longitudinal information from previous assessments could improve detection of changes in depressive symptoms over time—an important clinical use case for continuous monitoring. Threshold Validation: The clinical classification thresholds were assigned based on clinical judgment. Data-driven optimization against clinical outcomes should inform threshold refinement [2, 3].

    Multimodal Depression Detection An Integrated Multimodal Framework for Automated Depression Severity Assessment · 2026 · DOI
  • into it • Better accuracy: A hybrid of BERT stress and facial emotion recognition • Integration of the comprehensive data: Integrates both visual (facial expressions) and textual (stress questionnaires) data to have a fuller evaluation • Enhanced feature extraction: The need for more advanced techniques that extract comprehensive features from multimodal data sources • Scalable, long-term studies: There is a requirement for a larger and more varied dataset that encompasses participants over an extended period to enhance the generalizability of the findings System for Mental Health Prediction…

    Intelligent IoT-Based Mental Health Prediction Framework for Smart Cities Using Deep Learning: Integrating Facial Emotions and Questionnaire · 2026 · DOI
  • The paper reports class-wise performance variation across pain levels (BL1, PA1-PA4) with some pain categories achieving 75-87% precision while others reach 85%+, but does not analyze whether misclassification patterns correlate with specific pain intensity levels or physiological signal characteristics. Investigation into whether intermediate pain levels (PA2, PA3) are inherently more difficult to discriminate from fused signals would inform model improvements.

    Hybrid Deep Model for Pain Intensity Classification Using Fused ECG, EMG, and GSR Signals · 2026 · DOI
  • The multi-head attention mechanism in the hybrid model is applied without specification of the number of attention heads, attention dimension, or justification for these hyperparameters in the context of temporal pain signal modeling. Investigation of how attention head configurations (e.g., 4, 8, 16 heads) affect pain classification performance with fused ECG-EMG-GSR signals could optimize the attention mechanism design.

    Hybrid Deep Model for Pain Intensity Classification Using Fused ECG, EMG, and GSR Signals · 2026 · DOI
  • The paper emphasizes that optimal signal sequence selection achieves performance improvements 'without increasing computational costs' but provides no computational complexity analysis, inference time comparisons, or memory footprint measurements for the BiLSTM-MHAT-CNN hybrid model versus the CNN baseline. Runtime and hardware requirements must be quantified to support claims of practical applicability in real-time automated pain recognition systems.

    Hybrid Deep Model for Pain Intensity Classification Using Fused ECG, EMG, and GSR Signals · 2026 · DOI
  • The study evaluates pain intensity classification using only 5-fold cross-validation on a single proprietary dataset without reporting dataset size, participant demographics, pain induction methods, or signal sampling rates. External validation on publicly available pain-related physiological signal datasets (e.g., BioVid, UNBC-McMaster) is necessary to assess generalization of the hybrid model across different pain assessment protocols and populations.

    Hybrid Deep Model for Pain Intensity Classification Using Fused ECG, EMG, and GSR Signals · 2026 · DOI
  • The hybrid BiLSTM-MHAT-CNN model achieves 80.5% accuracy on the pain intensity dataset, substantially lower than the CNN baseline's 94.14%, yet the paper provides no ablation study isolating the contribution of BiLSTM, multi-head attention, and CNN components. Specific ablation experiments are needed to determine which architectural components degrade performance and whether attention mechanisms conflict with pain recognition from fused ECG, EMG, and GSR signals.

    Hybrid Deep Model for Pain Intensity Classification Using Fused ECG, EMG, and GSR Signals · 2026 · DOI
  • The study demonstrates that signal sequence ordering (GSR–ECG–EMG vs. ECG–GSR–EMG) significantly impacts hybrid BiLSTM-MHAT-CNN model performance, but lacks systematic investigation of all possible permutations and theoretical justification for optimal ordering. Future work should enumerate and evaluate all six permutations of the three physiological signals to establish evidence-based signal sequencing guidelines for multimodal pain intensity classification.

    Hybrid Deep Model for Pain Intensity Classification Using Fused ECG, EMG, and GSR Signals · 2026 · DOI
  • Although this study advances our understanding of MM and MA in vocal ERA, several limitations warrant consideration. First, we followed the discrete emotion framework, including six distinctive emotional expressions (Barrett, 2017). Future research should also incorporate additional approaches to facilitate the generalization of results and enable comparisons with previous cross-modal studies. Additionally, we believe that research would benefit from considering various types of stimuli, such as vocal bursts, words, and nonsense words, to better understand the nature of MA in emotion recognition across emotions (Lausen and Hammerschmidt, 2020).

    Metacognitive monitoring and calibration in the vocal emotion recognition task · 2026 · DOI
  • Lightweight attention-based architectures balancing predictive performance and computational efficiency have not been explored; future research should develop and evaluate streamlined attention variants (e.g., efficient Transformers, local attention) for practical deployment of audio-based multiclass depression screening.

    A Robust Hybrid Deep Learning Model for Multiclass Depression Classification from Speech Audio · 2026 · DOI
  • Dataset expansion must prioritize speaker-balanced and multilingual speech corpora to enable external benchmarking and cross-population validation; current evaluation is limited to a small single-language dataset, constraining generalizability of the hybrid deep learning multiclass depression classification model.

    A Robust Hybrid Deep Learning Model for Multiclass Depression Classification from Speech Audio · 2026 · DOI
  • Statistical significance testing and confidence interval estimation were not performed due to single train–test split and limited dataset size; rigorous cross-validation and statistical validation protocols must be implemented to establish confidence in the multiclass depression classification results.

    A Robust Hybrid Deep Learning Model for Multiclass Depression Classification from Speech Audio · 2026 · DOI
  • EEG data are available in the dataset repository but were not utilized; multimodal fusion integrating EEG with audio signals for improved robustness in multiclass depression severity classification remains unexplored and is explicitly positioned as a future research direction.

    A Robust Hybrid Deep Learning Model for Multiclass Depression Classification from Speech Audio · 2026 · DOI
  • Feature-level ablation study was not conducted during hybrid deep learning architecture benchmarking; systematic ablation of individual components (CNN, GRU, BiLSTM, Transformer) and feature contributions (MFCC, STFT, MelSpec) is needed to determine which elements drive multiclass depression classification performance.

    A Robust Hybrid Deep Learning Model for Multiclass Depression Classification from Speech Audio · 2026 · DOI

Most-cited papers in Emotion and Mood Recognition

Most recent work

Find a gap in your own Emotion and Mood Recognition sub-topic

This page shows what the Emotion and Mood Recognition literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.

Open the Research Gap Finder →

Related topics in Psychology

99 open questions have been extracted from the limitations and future-work passages of 519 Emotion and Mood Recognition papers in our library. Each one below links back to the study that raised it, so you can read the original claim in context.

Tools for your next paper

Compare the categoryHonest roundups of the AI research tools, ours listed alongside the alternatives.

Command palette

Jump anywhere, run any action.