Open research questions in Computational and Text Analysis Methods
254 unresolved questions extracted from the limitations and future-work sections of 927 Computational and Text Analysis Methods papers in our library. Each links back to the study that raised it.
What the literature leaves open
Consequently, future research should investigate annotation elicitation methods that allow participants to convey the important aspects of the submissions coherently and concisely in a repeatable fashion while facilitating scalable analysis. Large dataset analysis and insight extraction with adequate steering control per user goals is an open problem, documented previously [24]. While PLACES provides a list of concrete failure examples, it also poses an open question for model developers and safety researchers about the adequate model response in such situations.
Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South · 2026further studies should examine the relationship between metaknowledge and modeling practices in different contexts, - future research should investigate the role of content knowledge in the use of modeling practices, - the study suggests that examining the interplay between meta and content knowledge in modeling practices is a promising area for future research
Relating Metaknowledge and Modeling Practices in Connection to Content Knowledge about the Phenomenon · 2026 · DOIfurther investigation of the differences between LLMs and human behavior, - exploration of the applications of the Bayesian framework in other domains, - examination of the effects of fine-tuning on LLMs in different contexts
Disentangling interaction and bias effects in opinion dynamics of large language models · 2026 · DOIThe lack of a systematic framework to quantify and compare interaction and bias effects in LLM opinion dynamics. The need to disentangle genuine interaction from systematic biases in large language models. The limited understanding of how biases affect opinion dynamics in large language models.
Disentangling interaction and bias effects in opinion dynamics of large language models · 2026 · DOIYet this assumption is rarely studied directly, leaving us with scant knowledge about what people actually learn when accessing information about norms like human rights.
The Language Barrier to Human Rights Information: A Global Analysis of Google Search Results · 2026 · DOIIndeed, delegating judgment to such systems may affect the heuristics underlying evaluative processes, suggesting a shift from normative reasoning toward pattern-based approximation and raising open questions about the role of LLMs in evaluative processes.
Moreover, this work also considers other underexplored dangers of AI development for the environment and, hypothetically, for sentient AI.
Future research could explore the use of larger and more complex generative models. Future research could investigate the effects of different training data on model outputs. Future research could develop more sophisticated methods for sentiment analysis and framing detection.
Training Generative AI on Social Media Data: Implications and Outputs - A Worked Out Example · 2026 · DOIThe gap between how generative AI models are developed and what researchers can study. The lack of transparency and cautious interpretation in AI research.
Training Generative AI on Social Media Data: Implications and Outputs - A Worked Out Example · 2026 · DOIApplying the UMS to other historical moments and events. Refining the UMS protocol to improve its effectiveness and efficiency. Exploring the potential applications of the UMS in other fields, such as social sciences and natural sciences.
The Undirected Multidisciplinary Sweep: A General Method for Detecting Structural Blindness in Historical Records · 2026 · DOIThe gap between what is logically obvious and what is operationally executed in historical research. The lack of a systematic approach to detecting structural blindness in historical records.
The Undirected Multidisciplinary Sweep: A General Method for Detecting Structural Blindness in Historical Records · 2026 · DOIThe study only evaluates five Large Language Models. The corpus is limited to Indian, Singaporean, and Nigerian English varieties. The study does not investigate the impact of cultural ghosting on human subjects.
When AI Writes, Whose Voice Remains? Quantifying Cultural Marker Erasure Across World English Varieties in Large Language Models · 2026 · DOIInvestigating the impact of cultural ghosting on human subjects. Developing strategies for culturally-aware alignment. Evaluating the effectiveness of cultural-preservation prompts in different cultural contexts.
When AI Writes, Whose Voice Remains? Quantifying Cultural Marker Erasure Across World English Varieties in Large Language Models · 2026 · DOICultural biases in LLMs. Limited ability of existing evaluation frameworks to capture cultural inclusivity. The need for a benchmark to evaluate Global Cultural Inclusivity in LLMs.
WorldView-Bench: A Benchmark for Evaluating Global Cultural Perspectives in Large Language Models · 2026 · DOIExisting benchmarking frameworks fail to adequately capture cultural bias. The lack of a benchmark for evaluating Global Cultural Inclusivity in LLMs.
WorldView-Bench: A Benchmark for Evaluating Global Cultural Perspectives in Large Language Models · 2026 · DOIFuture research should focus on integrating thematic areas and strengthening research alliances. Future research should examine the application of AI in crisis management and the development of responsible AI frameworks.
Mapping the Nexus: A Bibliometric Analysis of the Application of Artificial Intelligence (AI) in Polycrisis Research (2019 – 2026) · 2026 · DOIThe research area of polycrisis is fragmented, with no systematic mapping of the role of AI. The study identifies a gap in the integration of thematic areas and the development of responsible AI frameworks.
Mapping the Nexus: A Bibliometric Analysis of the Application of Artificial Intelligence (AI) in Polycrisis Research (2019 – 2026) · 2026 · DOILongitudinal empirical studies are needed to further explore the integration of deep learning and digital literacy. Inclusive AI-driven pedagogical designs should be developed and evaluated. Further research is needed to address the research gap in using AI for fundamental research processes.
Deep Learning and Digital Literacy: A Systematic Literature Network and Bibliometric Review · 2026 · DOIA significant research gap exists in using AI for fundamental research processes compared to its dominance in instructional assessment. The study identifies a need for further research on the application of deep learning in digital literacy.
Deep Learning and Digital Literacy: A Systematic Literature Network and Bibliometric Review · 2026 · DOITopic modelling is often criticised for being overly descriptive and lacking theoretical integration. The study acknowledges the limitations of traditional topic modelling, including reliance on oversimplified assumptions and limited theoretical integration. The study also acknowledges the challenges related to reliability, validity, and multilingual inclusivity.
Going beyond description: multilingual topic modelling and theoretical integration in comparative media analysis · 2026 · DOIThe analysis of actual media content across nations has received comparatively less attention and has relied largely on manual methods. There is a need for more effective methods for analysing large-scale textual datasets in comparative media analysis. The study aims to address the limitations of traditional topic modelling and contribute to the development of more effective methods for comparative media analysis.
Going beyond description: multilingual topic modelling and theoretical integration in comparative media analysis · 2026 · DOIThe information gap between medical institutions and vaccine-hesitant populations. The lack of effective public health communication strategies to address individual queries and monitor population-level discourse.
Japanese-Language AI Agent System for Human Papillomavirus Vaccine Infoveillance and Public Communication: Development and Feasibility Evaluation · 2026 · DOIFuture research should investigate the causes of bias in LLMs and develop solutions to mitigate bias. Future research should examine the implications of bias in LLMs for various applications. Future research should develop more comprehensive and nuanced methods for evaluating bias in LLMs.
There is a lack of understanding of the limitations and potential risks of LLMs, particularly in regards to bias. There is a need to investigate the causes of bias in LLMs and to develop solutions to mitigate bias.
The treatment-timing rule is mechanically related to output. The event-study comparison is asymmetric in a way that mimics a causal effect. The design problem identified in this paper is a challenge for future research.
Comment on Scientific production in the era of large language models · 2026
Most-cited papers in Computational and Text Analysis Methods
- Reliability of Content Analysis: The Case of Nominal Scale Coding · Public Opinion Quarterly · 1955 · 1,360 citations
- Text as Data · Journal of Economic Literature · 2019 · 1,156 citations
- Extracting Policy Positions from Political Texts Using Words as Data · American Political Science Review · 2003 · 860 citations
- Collaborating With ChatGPT: Considering the Implications of Generative Artificial Intelligence for Journalism and Media Education · Journalism & Mass Communication Educator · 2023 · 709 citations
- Out of One, Many: Using Language Models to Simulate Human Samples · Political Analysis · 2023 · 666 citations
- An Evaluation of Amazon’s Mechanical Turk, Its Rapid Rise, and Its Effective Use · Perspectives on Psychological Science · 2018 · 641 citations
- ChatGPT and the rise of large language models: the new AI-driven infodemic threat in public health · Frontiers in Public Health · 2023 · 592 citations
- Automated Text Analysis for Consumer Research · Journal of Consumer Research · 2017 · 541 citations
- Scaling Policy Preferences from Coded Political Texts · Legislative Studies Quarterly · 2011 · 483 citations
- Introduction—Topic models: What they are and why they matter · Poetics · 2013 · 360 citations
Most recent work
- Canon Formation in the Age of AI: Metadata Packet for Disambiguation, Training-Layer Selection, and Retrocausal Reception (v1.1) · Zenodo (CERN European Organization for Nuclear Research) · 2026
- The Reverse Turing Test: A Three-Stage Protocol for Detecting AI-Mediation Signatures in Human Text and Their Propagation to Model Training (v1.2) · Zenodo (CERN European Organization for Nuclear Research) · 2026
- Thematic analysis with open-source generative AI and machine learning: a new method for inductive qualitative codebook development · Humanities and Social Sciences Communications · 2026
- The Vanishing Context: A Preliminary Content Analysis of Relational Context Loss in AI-Generated Bullet-Point Structuring of Public Subsidy Information Pages · Zenodo (CERN European Organization for Nuclear Research) · 2026
- Large language models can predict the results of social science experiments · Nature · 2026
- A reporting checklist for large language models in behavioural science · Nature Human Behaviour · 2026
- An Overview of Large Language Models for Statisticians · The American Statistician · 2026
- Computational hermeneutics: evaluating generative AI as a cultural technology · Frontiers in Artificial Intelligence · 2026
- The ideological orientation of academic social science research 1960–2024 · Theory and Society · 2026
- The Emerging Market for Intelligence: How Firms Buy and Sell AI · The Journal of Economic Perspectives · 2026
Find a gap in your own Computational and Text Analysis Methods sub-topic
This page shows what the Computational and Text Analysis Methods literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →