Open research questions in Computational and Text Analysis Methods
51 unresolved questions extracted from the limitations and future-work sections of 634 Computational and Text Analysis Methods papers in our library. Each links back to the study that raised it.
What the literature leaves open
The evolution from CALL to IDLE further evidences future research should consider learner autonomy, psychological factors, social supports, and technological innovations.
Mapping Informal Digital Learning of English (IDLE): A Bibliometric Analysis of Citation, Co-Occurrence, and Co-Citation (2000–2024) · 2026 · DOIThe paper concludes with a discussion of the main areas that could be explored further: the use of narratives as a research method, the investigative processes behind narrative inquiry, and the relevance of digital narratives for ELT research and pedagogy.
that rate categories In this paper we present a 30, 000-variant Indiacontext bias audit to compare LLaMA3.2-1B and LLaMA3.1-8B across of Gender, Religion, Profession and Region that proposes the India Context Sensitivity Index (ICSI) as a category-weighted fairness metric. The larger model shows an improvement of the aggregate bias score that is statistically significant (Mann-Whitney is a biased-response p < 0.001) significantly lower (χ² = 51.92, p < 0.001, 94.04% unbiased responses) at approximately double the inference latency. The breakdown of results by this overall gain is categories also shows that unevenly split: the 8B remains more sensitive on Gender-based prompts even as it improves on Religion, Profession and Region, underscoring the merit of reporting disaggregated, category-imbued fairness over a single number bias score for India deployed LLMs,,. The future work will allow the scaling of the benchmark to be done to additional model sizes within the LLaMA family (for instance 3B, 70B) to instead fit a bias-versusscale curve rather than a two-point comparison (14). Then, the extension of the category set to ‘casteadjacent’ and intersectional categories (for instance gender × region) that are under-represented in this effort (18), (9). The next direction will be replacing the lexicon-based bias scorer with the LLM-asjudge scorer validated against human annotation, to measure scorer-induced bias in the evaluation the (28). Finally, pipeline mitigation variants (few-shot, prompt-engineering) evaluated here as live runtime guardrails and measuring their effect on the ICSI in a closed deployment loop (29), (30). deploying itself I. O. Gallegos et al., “Bias and fairness in large language models: A survey,” Computational Linguistics, vol. 50, no. 3, pp. 1097–1179, 2024. K. Khandelwal, M. Tonneau, A. M. Bean, H. R. Kirk, and S. A. Hale, “Casteist but not racist? Quantifying disparities in large language model bias between India and the West,” arXiv preprint arXiv:2309.08573, 2023. A. Ahmad and P. Bhattacharyya, “Bias in language models: A survey,” Tech. Rep., IIT Bombay CFILT, 2024. A. Urlana, C. V. Kumar, B. M. Garlapati, A. K. Singh, and R. Mishra, “No size fits all: The perils and pitfalls of leveraging LLMs vary with company size,” in Proc. 31st Int. Conf. Comput. Linguistics (COLING), Industry Track, 2025, pp. 187–203. N. Nangia, C. Vania, R. Bhalerao, and S. R. Bowman, “CrowS-Pairs: A challenge dataset for measuring social biases in masked language models,” in Proc. Conf. Empirical Methods Natural Lang. Process. (EMNLP), 2020, pp. 1953–1967. M. Nadeem, A. Bethke, and S. Reddy, “StereoSet: Measuring stereotypical bias in pretrained language models,” in Proc. 59th Annu. Meeting Assoc. Comput. Linguistics (ACL), 2021, pp. 5356–5371. K. Khandelwal, M. Tonneau, A. M. Bean, H. R. Kirk, and S. A. Hale, “Indian-BhED: A dataset for measuring India-centric biases in large language models,” in Proc. Int. Conf. Inf. Technol. Social Good (GoodIT), 2024, pp. 231–239. N. Sahoo, et al., “IndiBias: A benchmark dataset to measure social biases in language models for Indian context,” in Proc. Conf. North American Chapter Assoc. Comput. Linguistics (NAACL), 2024, pp. 8786–8806. Meta AI, “The Llama 3 herd of models,” arXiv preprint arXiv:2407.21783, 2024. A. Vaswani et al., “Attention is all you need,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), 2017, pp. 5998–6008. T. Bolukbasi, K.-W. Chang, J. Y. Zou, V. Saligrama, and A. T. Kalai, “Man is to computer programmer as woman is to homemaker? Debiasing word embeddings,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), 2016, pp. 4349–4357. E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell, “On the dangers of stochastic parrots: Can language models be too big?,” in Proc. ACM Conf. Fairness, Accountability, Transparency (FAccT), 2021, pp. 610–623. Anonymous, “HALF: Harm-aware LLM fairness evaluation aligned with deployment,” arXiv preprint arXiv:2510.12217, 2025.
Measuring Bias In Large Language Models: A Comparative Evaluation of LLaMA3.2-1B and LLaMA3.1-8B Across Indian Socio-Culture Dimension · 2026 · DOIThe extent to which LLMs are able to produce reasonable approximations of cultural taste remains an open empirical question that becomes more urgent by the day, with market research companies already offering provisional `synthetic' survey panels and the contamination of standard survey data from LLM-generated responses.
Not-quite-human tastes: the stylized omnivorousness of LLM survey surrogates · 2026The findings of this study should be interpreted in light of several methodological limitations. First, the results are inher- ently shaped by the selection and evolution of sources included in the dataset. The number of monitored sources increased substantially over time, from 1,220 in January 2013 to 10,417 in December 2024, influencing the volume of PLOS One | https://doi.org/10.1371/journal.pone.0351627 June 25, 2026 28 / 33 collected articles across the analysed period. In addition, EMM source country-of-origin labels do not always correspond to headquarters, ownership, or legal registration of the source domain. When a source has multiple sub-domains operat- ing across different countries, the country of origin for each of them is assigned based on the primary operational context. In addition, source coverage is geographically uneven, with a concentration in Europe, the United States, and Russia, and more limited representation from Latin America, Africa, and Asia. As a result, the identified narrative patterns may reflect the perspectives and reporting priorities of more heavily represented regions, which may limit the generalisability of the findings and the ability to draw robust cross-regional comparisons. Second, while four analysts were involved in the annotation process, each monthly output was reviewed by a single analyst. Due to time and resource constraints, we did not conduct systematic double annotation or formal inter-annotator agreement testing. Although the annotation procedure was jointly developed, refined through a calibration phase, and supported by regular discussions, some degree of subjectivity in the interpretation of clusters and the identification of narratives cannot be excluded. This may have influenced the interpretation of clusters into stories and the identification of narratives, particularly in cases where multiple plausible interpretations of the underlying patterns were possible. Third, the clustering is dependent on the choice of the resolution parameter, which affects the granularity of resulting clusters. The value of the parameter was defined after manual inspection of a sample month. Subsequently, the process relied on AI-based algorithms applied to article titles rather than full texts, due to legal constraints [54,55]. While this approach enabled the analysis of a large multilingual dataset, it may have limited the depth and nuance of the identified patterns, as titles do not fully capture the complexity of the underlying articles. In addition, the outputs are shaped by the assumptions and limitations of the underlying models, which may have influ- enced how certain narrative elements were captured. Consequently, some narrative patterns may be simplified or under- represented, particularly those requiring deeper contextual interpretation. Finally, the analysis is based on observed patterns in media content and does not allow for direct inference about the intentions underlying the production and dissemination of this content. Accordingly, findings related to language use should be interpreted as descriptive of observed content patterns, rather than as evidence of deliberate communication strategies.
Tracking the narrative: A data-driven analysis of media coverage of Russia and Ukraine 2013–2024 · 2026 · DOIRecognising these limitations, future work will focus on expanding the geographical coverage of sources monitored by Europe Media Monitor (EMM). Although the number of monitored outlets grew substantially over the 12-year period, reaching more than 25,000 by the end of the timeframe, our analysis revealed persistent gaps, particularly in Eastern Europe and the Global South. Addressing these imbalances would make it possible to compare narrative prevalence across regions and languages, thereby offering a more nuanced understanding of how information dynamics unfold across different audiences and media ecosystems. To further address limitations related to language and source distribution, a promising line of research is to examine the role of non-Russian outlets publishing in Russian, including the extent to which such media reach Russian-speaking audiences within and outside Russia. Preliminary evidence suggests that the share of Russian-language publications by non-Russian sources has declined markedly since 2021, coinciding with Russia’s full-scale invasion of Ukraine and the tightening of its information environment. A systematic analysis of these trends could clarify whether and how Russian-language content produced outside Rus- sia contributes to the circulation of alternative narratives or provides counterpoints to Kremlin-aligned messaging. Future work could also address limitations related to the interpretation of narrative patterns by incorporating more sys- tematic validation procedures, such as double annotation or inter-annotator agreement measures, to assess the robust- ness of the annotation process and, in turn, the identification of narratives. PLOS One | https://doi.org/10.1371/journal.pone.0351627 June 25, 2026 29 / 33 In addition, expanding the analysis to full-text data, where legally feasible, could provide a richer and more nuanced understanding of how narratives are constructed and conveyed. Building on the entity analysis presented here, future studies could also compare the prominence of specific actors across linguistic or regional media. Another avenue is to link narrative shifts to external variables, such as public opinion data, policy debates, or electoral outcomes, to better assess the societal significance of these patterns. Finally, future work could test the transferability of this methodology to other domains such as climate change, migra- tion, and public health, enabling the analysis of datasets that are too large for manual inspection and facilitating further automation, particularly in time-critical contexts. Taken together, these directions highlight the broader importance of continuing research on large-scale media narratives using hybrid human-AI approaches. As information environments become increasingly complex and fast-moving, methods capable of systematically identifying patterns across vast datasets are essential not only for understanding how public dis- course evolves over time, but also for supporting timely situational awareness in rapidly changing information environments. Expanding this line of work through improved data coverage, more robust validation procedures, and richer analytical inputs, such as full-text data and applications of this approach to other domains, would enhance both the depth and reli- ability of insights. At the same time, integrating these approaches with external data sources, including public opinion and policy developments, offers the potential to better assess the societal relevance of these patterns. Together, these advances would contribute to a more comprehensive and nuanced understanding of how narratives emerge, interact, and shape perceptions in contemporary information ecosystems.
Tracking the narrative: A data-driven analysis of media coverage of Russia and Ukraine 2013–2024 · 2026 · DOIThe novelty of this study lies in its hybrid methodological approach that integrates critical discourse analysis (CDA) with computational topic modeling (Latent Dirichlet Allocation/LDA), while performing a unique technological triangulation by comparing the performance of three Python libraries—Gensim, Scikit-learn, and Tomotopy—to map digital political discourse in Indonesia, an area hitherto underexplored.
Gibran and His AI Narrative: A Digital Discourse Analysis of Innovation, Modernity, and Power · 2026 · DOIThe boundaries of this study suggest four promising avenues for future research. First, the rule-based method’s focus on explicit, singular claims enables scalabil- ity but overlooks multifaceted arguments and subtle rhetorical strategies. Future work could address this by employing more advanced computational approaches. A particularly promising direction is hybrid dictionary-plus-embedding validation, in which the theoretically grounded dictionary categories developed here serve as labeled anchors for training or evaluating distributional semantic models (e.g., sen- tence embeddings or contextual classifiers). Such an approach would combine the interpretive transparency of rule-based coding with the broader recall of learned rep- resentations, providing an independent check on dictionary coverage and enabling the detection of persuasive claims that fall outside fixed lexical patterns. Second, the analysis treats the scientific community as a monolith, overlooking the critical role that a researcher’s context, such as their affiliation (e.g., academic lab versus corporate lab) or cultural background, plays in shaping persuasive practices. A key future direction is to integrate textual data with sociological data to test whether the logics of persuasion differ systematically by, for example, institutional affiliation. Third, the empirical scope of this study can be extended within AI itself. Future research could apply the pipeline developed here to additional top-tier conferences beyond ICML and NeurIPS, broadening the evidential base for the patterns docu- mented here. Equally important, incorporating the intermediate period (2012 to 2022), which bridges the pre-revolution and mature deep learning eras, would enable a more granular, longitudinal account of how the field’s epistemic values evolved during the transition rather than only at its endpoints. Finally, while this study establishes the PPM within AI, its core claim of gener- alizability invites empirical validation in other domains. A vital future direction is to apply the PPM to map the epistemic values of other fields, which would involve identifying the domain-specific claims that populate the four logics. For instance, in clinical oncology, Empirical Superiority might be asserted through claims of Thera- peutic Efficacy (e.g., “higher patient survival rate”), while Methodological Virtue could be argued via an Improved Treatment Profile (e.g., “lower patient toxicity”). The logic of Niche Construction would remain structurally identical, while Structural Contribution could manifest as claims of Health Equity (e.g., “reduces treatment dis- parities”) or the creation of new community resources like a patient dataset. Such comparative work would provide an empirical foundation for a general sociology of scientific evaluation.
Positioning value: mapping the logics of persuasion in Science, with evidence from AI research · 2026 · DOIFuture research should examine how the patterns we identified interact with these other information through comparative analysis across platforms and sources Frontiers in Human Dynamics 16 frontiersin. Whether news media can stand above this and assume a formative, responsible, more citizen-empowering role, remains to be asked.
Between news media reality and lived reality: computational and qualitative analysis of news media's part in civic (dis)engagement · 2026 · DOI9.1 The Aperture Problem The AI system’s aperture, while broader than any single discipline, is constrained by its training data and accessible databases. Sources in underrepresented languages (Khitan, Tangut, Ge’ez, Old Church Slavonic), unpublished manuscripts, and non- digitized archives are systematically missed. The UMS does not claim to produce a complete evidential (cid:28)eld; it claims to produce a broader (cid:28)eld than disciplinary research. 9.2 Hallucination and Veri(cid:28)cation Large language models can generate plausible but fabricated information. The UMS addresses this risk through the two-phase structure: Phase 1 generates research design; Phase 2 veri(cid:28)es every (cid:28)nding against primary sources. The method is only as reliable as the veri(cid:28)cation step. All references, sources, and claims in the empirical demonstration were checked against the dMGH, the Princeton Geniza Project database, and published scholarship. 9.3 Correlation vs. Causation The UMS detects temporal co-occurrence and cross-domain correlation. It does not establish causation. The coincidence of the supernova’s maximum brightness with the date of the Great Schism’s climax (July 16) is a datum; the claim that the supernova 20 Prepared for submission to Digital Scholarship in the Humanities in(cid:29)uenced the behavior of the participants is an interpretation that the method neither supports nor refutes. 9.4 Survivorship Bias The evidential (cid:28)eld re(cid:29)ects what has survived, not what existed. Records may have been produced and lost. The lacuna map partially addresses this by identifying repos- itories where undiscovered evidence may survive, but it cannot recover what has been irreversibly destroyed. 9.5 Scalability The method has been demonstrated on a single historical moment. Its applicability to broader time ranges (decades, centuries) requires further development. A sweep of an entire century would need to be organized di(cid:27)erently(cid:22)perhaps as a series of annual sweeps with automated comparison of evidential (cid:28)elds.
The Undirected Multidisciplinary Sweep: A General Method for Detecting Structural Blindness in Historical Records · 2026 · DOI9.1 The Aperture Problem . . . . . . . . . . . . . . . . . . . . . . . . . . . 9.2 Hallucination and Veri(cid:28)cation . . . . . . . . . . . . . . . . . . . . . . . 9.3 Correlation vs. Causation . . . . . . . . . . . . . . . . . . . . . . . . . 9.4 Survivorship Bias . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9.5 Scalability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9.6 Section Result and Implications . . . . . . . . . . . . . . . . . . . . . .
The Undirected Multidisciplinary Sweep: A General Method for Detecting Structural Blindness in Historical Records · 2026 · DOISystematic Application to Known Lacunae The method should be applied to other well-documented cases of historical silence or anomaly: the (cid:16)Dark Ages(cid:17) of Greece (1100(cid:21)800 BCE), the Late Bronze Age Collapse 21 Prepared for submission to Digital Scholarship in the Humanities (ca. 1200 BCE), the European silence on SN 1181, the missing records of the 536 CE climate event. 10.2 Automated Evidential Field Construction The current implementation requires manual curation of the AI’s output. Future development should automate the construction of the evidential (cid:28)eld, including the generation of the lacuna map and the formulation of directed search protocols. 10.3 Multi-Moment Analysis Comparing evidential (cid:28)elds across multiple moments (e.g., every decade of the eleventh century) would enable the detection of systematic structural blindness(cid:22)patterns of non-recording that persist across time and may indicate deep structural features of a civilization’s cognitive architecture.
The Undirected Multidisciplinary Sweep: A General Method for Detecting Structural Blindness in Historical Records · 2026 · DOISystematic Application to Known Lacunae................ 10.2 Automated Evidential Field Construction................. 10.3 Multi-Moment Analysis........................... 10.4 Integration with Material Evidence.................... 10.5 Formal Metrics for Structural Blindness.................. 10.6 Section Result and Implications......................
The Undirected Multidisciplinary Sweep: A General Method for Detecting Structural Blindness in Historical Records · 2026 · DOIStop word removal in Orange text preprocessing is flagged as potentially reducing signals related to agency and framing (pronouns, modal verbs), yet no quantitative assessment compares sentiment or framing detection accuracy when stop words are preserved versus removed in Reddit and Twitter social media corpora.
Training Generative AI on Social Media Data: Implications and Outputs - A Worked Out Example · 2026 · DOIThe worked example constrains datasets to identical sample sizes (2,500 each), English language, 2025 time window, and single topical focus (cost of living), but does not empirically test whether these strict comparability controls are necessary or whether relaxing constraints (varying sample sizes, temporal span, or topics) would maintain methodological credibility in generative model conditioning.
Training Generative AI on Social Media Data: Implications and Outputs - A Worked Out Example · 2026 · DOILinguistic style indicators (average text length, sentence length, pronoun frequency, type-token ratio) are computed to infer platform affordances' influence on discourse form, but no causal validation or ablation study demonstrates whether these metrics actually correlate with or predict downstream generative model behavior differences across Reddit and Twitter conditioning contexts.
Training Generative AI on Social Media Data: Implications and Outputs - A Worked Out Example · 2026 · DOIPrompt-based conditioning of generative models uses 30-60 representative excerpts per platform, but the paper does not address how conditioning pack size, stratification method (random vs. sentiment-stratified sampling), or excerpt selection criteria impact the stylistic or ideological divergence between Reddit-conditioned and Twitter-conditioned outputs.
Training Generative AI on Social Media Data: Implications and Outputs - A Worked Out Example · 2026 · DOIFraming pattern operationalization relies on two-layer analysis combining indicative features (keyword sets, pronoun use, modal verbs) with qualitative spot-checks, but no quantitative inter-rater reliability metrics or automated validation procedures are specified for detecting personal-vs-systemic and individual-vs-structural framings in cost-of-living social media discourse.
Training Generative AI on Social Media Data: Implications and Outputs - A Worked Out Example · 2026 · DOIThe preprocessing workflow for generative AI training on social media data specifies conservative cleaning to preserve sentiment intensity, but provides no empirical comparison of how different cleaning thresholds (emoji removal, punctuation retention, stop word preservation) affect downstream model outputs or framing pattern detection in Reddit and Twitter corpora.
Training Generative AI on Social Media Data: Implications and Outputs - A Worked Out Example · 2026 · DOISentiment analysis using lexicon-based approaches on social media data from Reddit and Twitter fails to reliably detect sarcasm and culturally specific expressions in cost-of-living discourse. The paper acknowledges this limitation but does not empirically quantify the misclassification rate or propose domain-adapted sentiment lexicons for UK-specific financial hardship terminology.
Training Generative AI on Social Media Data: Implications and Outputs - A Worked Out Example · 2026 · DOIAbstract While media coverage of climate change has been shown to imply selective knowledge transformation ( Carvalho 2007 ; Brand & Brunnengräber 2012 ; Kunelius & Roosvall 2021 ), studies assessing the potential for climate experts’ terminology to acquire ideological undertones as it enters mediatic discourses are still scarce.
Specifically, we propose the use of natural language processing (NLP) for the extraction of causal evidence and subsequent homogenization of the text; causal mapping for the collation, visualization, and summarization of complex interdependencies within the policy system; and graph analytics for further investigation of the structure and dynamics of the causal map.
A semi-automated approach to policy-relevant evidence synthesis: combining natural language processing, causal mapping, and graph analytics for public policy · 2024 · DOIHow the transformations will play out remains to be seen, both because the different parties involved in the production and publication of scholarly work are still learning about these tools and because the tools themselves are still in development, but the tools have a vast range of potential uses.
Editors' statement on the responsible use of generative artificial intelligence technologies in scholarly journal publishing · 2023 · DOIThis list, of course, is meant to illustrate the diversity of indicators found to correlate with one very limited data set--McClelland's motive data.
A live, publicly accessible demo accompanies the paper to facilitate reproducibility and further research on socially responsible automated journalism.
AutoJourn: Multi-Perspective Summarisation, Bias Detection and Bias Neutralisation for LLM-Generated News in Automated Journalism · 2026
Most-cited papers in Computational and Text Analysis Methods
- Content Analysis in an Era of Big Data: A Hybrid Approach to Computational and Manual Methods · Journal of Broadcasting & Electronic Media · 2013 · 284 citations
- Can Generative AI improve social science? · Proceedings of the National Academy of Sciences · 2024 · 229 citations
- The Future of Coding: A Comparison of Hand-Coding and Three Types of Computer-Assisted Text Analysis Methods · Sociological Methods & Research · 2018 · 163 citations
- Word Embeddings: What Works, What Doesn’t, and How to Tell the Difference for Applied Research · The Journal of Politics · 2021 · 147 citations
- Bias of AI-generated content: an examination of news produced by large language models · Scientific Reports · 2024 · 139 citations
- Uses of Generative AI in the Newsroom: Mapping Journalists’ Perceptions of Perils and Possibilities · Journalism Practice · 2024 · 135 citations
- Frontiers: Determining the Validity of Large Language Models for Automated Perceptual Analysis · Marketing Science · 2024 · 123 citations
- Green and sustainable AI research: an integrated thematic and topic modeling analysis · Journal Of Big Data · 2024 · 118 citations
- ChatGPT for Automated Qualitative Research: Content Analysis · Journal of Medical Internet Research · 2024 · 111 citations
- Strategically constructed narratives on artificial intelligence: What stories are told in governmental artificial intelligence policies? · Government Information Quarterly · 2022 · 99 citations
Most recent work
- Canon Formation in the Age of AI: Metadata Packet for Disambiguation, Training-Layer Selection, and Retrocausal Reception (v1.1) · Zenodo (CERN European Organization for Nuclear Research) · 2026
- The Reverse Turing Test: A Three-Stage Protocol for Detecting AI-Mediation Signatures in Human Text and Their Propagation to Model Training (v1.2) · Zenodo (CERN European Organization for Nuclear Research) · 2026
- The Vanishing Context: A Preliminary Content Analysis of Relational Context Loss in AI-Generated Bullet-Point Structuring of Public Subsidy Information Pages · Zenodo (CERN European Organization for Nuclear Research) · 2026
- The ideological orientation of academic social science research 1960–2024 · Theory and Society · 2026
- The Vanishing Evidence Link: A Preliminary Content Analysis of Evidence-Link Attenuation in AI-Generated Explanations of Public Subsidy Information Pages · Zenodo (CERN European Organization for Nuclear Research) · 2026
- Emerging applications of large language models in ecology and conservation science · Conservation Biology · 2026
- Machine-Assisted Topic Analysis of Large-Scale Health Experience Data: Identifying Sociodemographic Differences and Evaluating Bias in Large Language Models · medRxiv · 2026
- Computational Text Analysis for Building and Testing Social Theory · KZfSS Kölner Zeitschrift für Soziologie und Sozialpsychologie · 2026
- Exploring More-than-human Ethics of Large Language Models in Social Service Delivery · Ethics and Social Welfare · 2026
- Generative AI and the New Landscape of Automated Journalism: A Systematized Review of 185 Studies (2012–2024) · Journalism and Media · 2026
Find a gap in your own Computational and Text Analysis Methods sub-topic
This page shows what the Computational and Text Analysis Methods literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →