Psychology · Research topic

Open research questions in Mental Health via Writing

61 unresolved questions extracted from the limitations and future-work sections of 291 Mental Health via Writing papers in our library. Each links back to the study that raised it.

What the literature leaves open

  • Conclusion Zero-shot LLMs’ performance was insufficient for clinical application in predicting PTSD symptoms from semi-structured interview transcripts.

    Zero-Shot Large Language Models for Preliminary Prediction of PTSD Symptoms From Clinical Interview Transcripts: Grands modèles de langage sans exemple pour la prédiction préliminaire des symptômes de TSPT à partir de transcriptions d'entrevues cliniques · 2026 · DOI
  • Large language models (LLMs) have shown promise in extracting clinically relevant information from unstructured language, but their ability to infer item-level PTSD symptom severity from clinical interviews remains unclear.

    Zero-Shot Large Language Models for Preliminary Prediction of PTSD Symptoms From Clinical Interview Transcripts: Grands modèles de langage sans exemple pour la prédiction préliminaire des symptômes de TSPT à partir de transcriptions d'entrevues cliniques · 2026 · DOI
  • Background: Depression frequently co-occurs with ADHD and autism spectrum disorder (ASD), but population-level differences in symptom expression between these groups remain underexplored.

    Population-Level Profiling of DSM-5 Depressive Symptoms Among Self-Reported ADHD and ASD Users on Twitter: An Exploratory Study Using Advanced NLP and Statistical Analysis · 2026
  • Several limitations should be noted. First, recall was assessed at the SRP level only; per-substance recall (the rate at which the classifier misses specific substance types among SRP-present records) was not independently evaluated. Second, validation samples of 100 cases per category provide reasonable but not definitive estimates; rare edge cases may be underrepresented. Third, these results are based on a single state's records and may not generalize to jurisdictions with different documentation practices, terminology, or populations.

    Validation of a Small Language Model for DSM-5 Substance Category Classification in Child Welfare Records · 2026 · DOI
  • Our findings provide novel evidence that embedding models optimized for semantic similarity represent psychopathological constructs with variable fidelity. Models demonstrate some ability to consistently identify broad thematic domains, but cross-model inconsistencies at finer levels of granularity suggest representations may not yet capture the hierarchical, interrelated nature of psychopathology. Care must therefore be taken when working with cosine similarities derived from different LLMs. Future research should validate LLM-derived structures against empirically validated models of psychopathology and develop strategies to mitigate issues like range restriction. A critical next step is identifying which embedding models recover empirically-derived factor structures, and what model characteristics distinguish strong from weak performers. Prior work has demonstrated that some models can accurately predict factor structure (Kambeitz et al., 2025; McElroy et al., 2024; Ringwald et al., in press); systematic comparison could clarify which models are most suitable for structural analysis. Research should also examine whether findings from embedding models extend to generative LLM outputs, assessing whether generative processes amplify, distort, or remedy these consistency issues. Until targeted improvements address these issues, embedding-based tools for psychopathology research require careful evaluation before application.

    Representational Idiosyncrasy of Psychopathology Across Text Embedding Models: Implications and Practical Suggestions for Clinical Psychological Science · 2026 · DOI
  • Several limitations should be acknowledged. First, the closed-source nature of many models limits conclusions about the precise impact of architectural and training influences on representational structure. Second, our findings reflect relative similarity between model-derived solutions, not absolute fidelity to a 'ground truth' of psychopathology. Third, fundamental questions remain about what cosine similarities represent and how they relate to covariance structures (Steck et al., 2024); the present study did not attempt to align cosine similarities with empirical correlation structures through transformation, though such methods exist (Li et al., 2020; Mu et al., 2017; Su et al., 2021) and may partially resolve the inconsistencies documented here. Relatedly, the arbitrary orientation of embedding spaces complicates direct cross-model comparison; alignment procedures may address this (Dev et al., 2021), but in their absence it remains unclear to what extent observed structural differences reflect true representational divergence versus geometric residual variance. Fourth, we did not evaluate whether embedding- derived similarities predict empirical inter-item correlations. This leaves open the question of practical utility. Fifth, the IPIP-NEO-300 and IPIP-HEXACO archival files lack participant demographics (e.g., race/ethnicity, culture, sex/gender, SES), limiting generalizability of the empirical benchmarks. Finally, and perhaps most importantly, generative LLMs were outside the REPRESENTATIONAL IDIOSYNCRASY IN LLMS 37 scope of this investigation, yet end users in clinical and research contexts will predominantly interact with generative models. Whether the representational inconsistencies identified here extend to or are remediated by generative processes remains a critical open question for future work.

    Representational Idiosyncrasy of Psychopathology Across Text Embedding Models: Implications and Practical Suggestions for Clinical Psychological Science · 2026 · DOI
  • This study applied RNN to predict depressive symptoms among college students, representing an active exploratory effort in this field. However, there are several limitations to consider. First, all participants were drawn from a single university, which may limit generalizability. Institutional and regional factors - such as social economic context, access to mental health resources, academic stress, and cultural attitudes toward mental illness - may have influenced both the prevalence and predictors of depressive symptoms. Future studies should replicate and validate these findings across diverse institutions and cultural settings. Second, participant attrition over four years was relatively high, a common challenge in longitudinal research. Although no significant differences in depressive symptom prevalence were observed between participants ARTICLE IN PRESS ARTICLE IN PRESS ACCEPTED MANUSCRIPT retained and those lost to follow-up (see Supplementary Table S8), potential residual bias cannot be completely ruled out. Last, while a standard RNN with ReLU activation was selected for its computational efficiency and lower parameter complexity—which reduce the risk of overfitting given our sample size—this architecture is inherently susceptible to the vanishing gradient problem. This may limit its capacity to capture very long-range temporal dependencies compared to gated architectures such as LSTM or GRU. Although the moderate length of temporal sequences in our dataset reduces the potential impact of this limitation, we acknowledge that alternative architectures could yield different results. Future studies with extended follow-up periods or larger sample sizes may benefit from exploring LSTM, GRU, or other advanced temporal modeling approaches to further validate and extend our findings.

    Predicting depressive symptoms among Chinese college students using recurrent neural networks with longitudinal data · 2026 · DOI
  • Emotional Ambiguity: The system may misclassify complex or mixed emotions, resulting in less relevant stories. 2. Lack of Voice Integration: Current implementation is text-based; no voice narration or speechbased input has been added. 3. Limited Dataset: The model performance depends on the emotional dataset; limited emotional diversity may affect output richness. 4. Subjectivity in Evaluation: Emotional relevance can vary among users due to individual psychological differences.

    AI-Based Emotional Story Generator for Therapy · 2026 · DOI
  • This study had small sample size due to limited availability of resources like only one investigator, time constraints of participants etc. Thus, authors are aware that small sample size of this study is likely to be underpowered. Snowball sampling helped in identifying participants willing to participate in the study but limited the representativeness and generalizability of the findings. A brief, experimentally induced meaning-making manipulation was involved, therefore to avoid priming, practice and recency effects baseline assessment of study variables could not be conducted. Experimental and control tasks differed in duration which may have created expectancy effects or differences in participant engagement. This should be considered while interpreting the findings. For future studies, implementing an active control condition matched for duration and engagement such as neutral expressive writing about daily activities can help to control expectancy effects. It will also help to disentangle specific effects of meaning making from non-specific factors. Although, mediation analyses were performed to examine theoretically drawn pathways, the mediator and outcome variables were assessed at post-test only. Therefore, temporal ordering among mediator and outcome variables cannot be established. Consequently, mediation results should be interpreted as model consistent associations rather than definitive evidence of causal relationship. The effects observed for memory outcomes should be interpreted with caution. Along with modest magnitude of effects, the sample was of young adults who show relatively stable cognitive functioning, limiting the difference in memory performance. The PGI memory scale is an older test, mainly developed for neuropsychological assessment. As such, it may be less sensitive to subtle experimental effects in non-clinical population. Therefore, future research could include more contemporary memory measures like Wechsler Memory Scale or n-back working memory tasks to precisely catch subtle cognitive changes. This study did not assess other important factors influencing affect and cognitive functions, such as personality traits, intelligence, or coping styles. Including these variables in future research would offer a more comprehensive understanding of their interrelations. Additionally, the stress induced was brief and of low intensity; stronger or prolonged stressors may provide clearer insights into the effects of stress on affect and cognition. Although subjective stress ratings were taken, stress was not the primary target of manipulation, but the reliance on self-report limits the strength of conclusions regarding stress induction. Considering physiological and behavioural measures will provide more comprehensive assessment of stress-related responses. The study assessed only selected components of cognitive functioning, whereas cognition is multifaceted. Examining additional cognitive domains in relation to stress, meaning, and resilience would be valuable. As outcomes of the study were assessed using self-report measures and the participants were briefed about the nature of the study, demand characteristics and expectancy effects can’t be ruled out.

    Meaning Making Under Stress: The Mediating Role of Resilience in Affect and Memory · 2026 · DOI
  • Future work may explore applications to social media data for population- level monitoring, a larger dataset for suicide detection at clinical hospital stay level, extension to other mental health conditions such as depression and anxiety disorders, and real-world clinical deployment validation. approach evidence extraction to support interpretability, this system has not been validated by clinicians in real-world settings.

    Enhancing Suicide Risk Classification: A Multi-Stage Framework with Sentence-Level Waterfall Architecture for Clinical Notes Analysis · 2026 · DOI
  • Alternative architectures, regularization algorithms, and optimization techniques could increase deep learning performance and should be investigated in the future. distinct feature extraction approaches or domain-specific embed- dings may provide distinct performance patterns, which should be explored in future investigations. The major limitations are that despite the optimistic outcomes, this study has numer- ous limitations that must be acknowledged.

    Social media–driven machine learning approach to identify emotional support needs in cancer care · 2026 · DOI
  • The generalizability of BERT-based schizophrenia classification to populations with varying literacy levels, cultural discourse norms, or non-English languages has not been tested. Cross-linguistic validation and evaluation across diverse demographic populations with different mental-health vocabulary conventions are needed to assess whether morphological and structural markers remain invariant across linguistic and cultural contexts.

    On the interface between linguistics, computer science and psychiatry: analyzing textual key-factors affecting BERT-based classification of schizophrenia in social media texts · 2026 · DOI
  • Topicality effects on semantic coherence salience in schizophrenia detection remain theoretical; the paper cannot yet determine whether BERT captures deeper grammatical or coherence disruption patterns related to topic versus simply capturing differences in linguistic information density and discourse register across genres. Cross-topic coherence analysis with controlled semantic and syntactic complexity is required.

    On the interface between linguistics, computer science and psychiatry: analyzing textual key-factors affecting BERT-based classification of schizophrenia in social media texts · 2026 · DOI
  • The interaction between text length, discourse genre, and schizophrenia linguistic markers appears additive rather than interactive, but this relationship has not been formally tested across controlled genre conditions with minimum-length thresholds. Systematic manipulation of both text length and topic/genre type is needed to establish genre-specific minimum-length requirements for reliable BERT-based classification.

    On the interface between linguistics, computer science and psychiatry: analyzing textual key-factors affecting BERT-based classification of schizophrenia in social media texts · 2026 · DOI
  • The r/AskDocs subreddit analysis revealed lexical overfitting with only 18 samples; this requires expansion to larger samples across multiple health-adjacent social media contexts to quantify the degree to which explicit disorder-related vocabulary inflates classification performance and to establish minimum dataset sizes needed to control lexical bias in mental-health NLP models.

    On the interface between linguistics, computer science and psychiatry: analyzing textual key-factors affecting BERT-based classification of schizophrenia in social media texts · 2026 · DOI
  • Morphological and phonotactic errors are hypothesized as critical grammatical markers of schizophrenia based on procedural memory impairment, but the paper does not systematically extract or weight these morphological features in the BERT classification pipeline. A dedicated feature engineering approach should quantify and preserve morphological error patterns while filtering semantic mental-health vocabulary.

    On the interface between linguistics, computer science and psychiatry: analyzing textual key-factors affecting BERT-based classification of schizophrenia in social media texts · 2026 · DOI
  • The specific relationship between grammatical structure, coherence relations, and the attention mechanisms allocated by BERT to these linguistic structures remains unexamined. Future work should conduct attention visualization analyses to determine which grammatical markers and coherence features the model prioritizes when classifying schizophrenia in social media texts.

    On the interface between linguistics, computer science and psychiatry: analyzing textual key-factors affecting BERT-based classification of schizophrenia in social media texts · 2026 · DOI
  • The current system performs snapshot sentiment analysis on individual tweets without temporal modeling; the paper explicitly identifies leveraging temporal analysis to model long-term mental health behavioral patterns and longitudinal trends in mental health discourse on social media as an unaddressed gap.

    A Machine Learning-Based Sentiment Analysis Framework for Twitter Tweets Using Natural Language Processing · 2026 · DOI
  • The sentiment analysis framework is currently validated only as a standalone interactive web application; the paper identifies implementing the proposed system within actual healthcare and counseling frameworks as a critical gap requiring integration testing, clinical validation, and assessment of real-world deployment in professional mental health settings.

    A Machine Learning-Based Sentiment Analysis Framework for Twitter Tweets Using Natural Language Processing · 2026 · DOI
  • The proposed system uses binary classification (concern detected vs. no concern) for mental health issues on Twitter; the paper specifies the need to adapt the approach toward multi-class classification to detect and differentiate between other specific mental health problems (depression, anxiety, bipolar disorder, etc.) rather than generic concern detection.

    A Machine Learning-Based Sentiment Analysis Framework for Twitter Tweets Using Natural Language Processing · 2026 · DOI
  • The sentiment analysis system is currently trained and validated only on English-language Twitter tweets; the paper identifies extending multilingual support to other languages as a concrete gap, requiring adaptation of the NLP preprocessing pipeline and VADER sentiment analyzer for non-English social media discourse in mental health contexts.

    A Machine Learning-Based Sentiment Analysis Framework for Twitter Tweets Using Natural Language Processing · 2026 · DOI
  • The current Logistic Regression model with TF-IDF features achieves 82-85% accuracy on mental health concern detection in Twitter tweets; the paper explicitly identifies exploring alternative deep learning architectures (such as LSTM, CNN, or transformer-based models) as necessary to boost predictive performance beyond current baseline levels.

    A Machine Learning-Based Sentiment Analysis Framework for Twitter Tweets Using Natural Language Processing · 2026 · DOI
  • The ablation study demonstrated that negation handling and intensifier modeling are critical preprocessing components (5.1% and 1.7% performance impact respectively), but the study does not specify which negation patterns or intensifier types are most predictive of depression in English versus Arabic. Future work should conduct language-specific linguistic analysis to identify the particular negation constructions and intensifier categories that differentiate depressive from non-depressive social media discourse in each language.

    Multilingual depression screening via social media: comparative analysis of machine learning models on English and Arabic text · 2026 · DOI
  • The current framework processes only text-based features from social media posts for depression detection. Future work should integrate multimodal features including images alongside text and develop longitudinal models capable of capturing behavioral trends over time to provide deeper insights into mental health trajectories and temporal patterns of depressive expression.

    Multilingual depression screening via social media: comparative analysis of machine learning models on English and Arabic text · 2026 · DOI
  • The multilingual depression screening framework was exclusively trained and validated on Twitter (X) data, limiting generalizability across social media platforms and user demographics. The study should expand evaluation to include data from Facebook, Instagram, and Reddit to assess whether the RBF-SVM classifier with TF-IDF preprocessing maintains performance consistency across different platform-specific linguistic styles and depression expression conventions.

    Multilingual depression screening via social media: comparative analysis of machine learning models on English and Arabic text · 2026 · DOI

Most-cited papers in Mental Health via Writing

Most recent work

Find a gap in your own Mental Health via Writing sub-topic

This page shows what the Mental Health via Writing literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.

Open the Research Gap Finder →

Related topics in Psychology

61 open questions have been extracted from the limitations and future-work passages of 291 Mental Health via Writing papers in our library. Each one below links back to the study that raised it, so you can read the original claim in context.

Tools for your next paper

Compare the categoryHonest roundups of the AI research tools, ours listed alongside the alternatives.

Command palette

Jump anywhere, run any action.