Open research questions in Subtitles and Audiovisual Media
29 unresolved questions extracted from the limitations and future-work sections of 896 Subtitles and Audiovisual Media papers in our library. Each links back to the study that raised it.
What the literature leaves open
However, free online corpora of video game dialogue in languages besides English and Japanese (such as the “FIGS” languages—French, Italian, German, and Spanish) are lacking.
A multi-language video game dialogue corpus · 2026For Learners: Because watching movies frequently with English subtitles is strongly linked to better TOEIC listening performance, students should transition away from Thai subtitles. Watching English-language films three to four times a week is recommended for steady improvement in listening skills. For Teachers: Instead of simply telling students to watch English movies, instructors should actively promote the use of English subtitles and create post-viewing activities, like dictation or comprehension quizzes, to solidify learning. Since Netflix is highly popular among students, teachers could also curate a list of shows on the platform that align with their students' skill levels. For Future Researchers: Upcoming studies should explore the motivations driving students' subtitle preferences—specifically, whether they use them out of habit or a deliberate desire to learn—as this context is crucial for interpreting data. Additionally, employing longitudinal methods and more diverse participant groups would help make the findings more universally applicable.
The current study is subject to several limitations. First, it only evaluated vocabulary learning at the level of form recall. To gain a more comprehensive picture of how textual enhancement in captions affects incidental vocabulary learning, future research can use several tests to measure vocabulary development at different sensitivity levels (form recognition, meaning recall/recognition). Second, a transferappropriate processing concern arises from the present design. According to the transfer-appropriate processing framework (Morris et al., 1977), learning performance is influenced by the congruency between the conditions of learning and the conditions of testing. Since learners in the captioned groups had access to written target forms during treatment while the uncaptioned group did not, the written posttest format may have systematically advantaged the captioned groups due to learning-testing congruency, making it difficult to determine the extent to which group differences reflected vocabulary learning rather than test-format effects. Third, the initial letters of the target items were supplied as clues to prevent learners from writing alternative items. This differs from L2 users’ production of vocabulary items in real-life situations. Fourth, the study was conducted at a single site with English-major university students, limiting generalisability. Vitta et al. (2022) highlighted the importance of multisite designs in L2 vocabulary research. Fifth, English-major students represent a specialised learner profile with potentially different characteristics than non-major learners. Future multisite research with more diverse learner profiles is therefore warranted. Sixth, the sample size per condition was relatively small. However, it should be mentioned that previous research on the effect of mode of input on incidental vocabulary learning has used the same number of participants, at least in one of their experimental conditions (Feng & Webb, 2020; Majuddin et al., 2021). Seventh, word- and learner-related variables, such as frequency, working memory, and visual support (see Teng & Cui, 2025; Kaderoğlu, 2026), were not investigated. Finally, to obtain a more fine-grained picture of the effect of the treatment conditions, future studies are recommended to include a test-only control group.
The Effect of Textual Enhancement on Incidental Learning of Single Words and Multi-word Items From L2 Captioned Viewing · 2026 · DOIHowever, with no concurrent studies to cope up and update this recent convenience of the TTS, little is known about its scale and forms of impact on the population of the Arabic speaking EFL learners.
Revisiting the Text to Speech Tool in EFL Learning in the Light of the Recent Rise in the Accessibility of the Technology · 2026 · DOIThis study had a number of limitations. It included only immediate and receptive learning posttests of form and form-meaning mapping, which limits any claims we can make regarding durability of learning, and did not include a test of word production. Additionally, the type of mapping selected (new L2 form to existing L1 concept) was specific within the broad possibilities for form-meaning mappings in L2 vocabulary learning (see Kida & Barcroft, 2018). As such, generalizable claims regarding learning are limited to L2 form-form and form-meaning mapping. One anonymous reviewer noted that we included both general and specific distractor types on the meaning recognition outcome test, tailored to the meaning associated with that pseudoword item, which could have impacted item responses; 76% of meaning recognition items (19/25) included only specific distractors, 12% (3/25) included only general distractors, and 12% (3/25) had one of each. Descriptive scores indicated notable differences in accuracy between items with specific (58% accurate), generic (67% accurate), or mixed (38% accurate) distractors, so we re-ran the main outcome analysis for the meaning recognition measure with distractor type (specific, generic, mixed) as a categorical predictor. In this analysis, distractor type was not a significant predictor of meaning recognition when other variables were included (b = (cid:1).41, SE =.27, p =.127), and did not change other results of the model, so the model without distractor type was retained as the best-fitting model. However, future studies utilizing similar outcome presentation modalities and examining learning of only specific or general items can better control for this potential source of variability in learning. The study was highly controlled, with pseudoword targets rather than real words, while the audio rate was standardized rather than modified to follow reading speed more closely at the individual level. Reading ahead was defined dichotomously yes/no in the analytic models, so it was not possible to examine the relative effects of lagging behind the audio when reading. Since our RQs focused only on the potential benefits of reading ahead, we did not explore differential impacts of reading precisely synchronously or lagging behind the audio; future studies are warranted which do so, thereby increasing precision and allowing for more specific theoretical predictions for instances when the eye and ear are synchronous, and when the eye lags behind the ear during RWL. The L1s among participants were diverse (see Table 2), which did not allow for an examination of the effects by L1, while recent findings from L2 English reading report evidence that L1-L2 differences may result in variable reading behavior, demonstrated by eye movements (Kuperman et al., 2025). This would be a fruitful area for future research. As the present study provides initial evidence that reading ahead of the audio during RWL facilitates faster word form recognition across instances and mapping form to meaning during contextualized learning of new L2 vocabulary, compared with reading synchronously or behind the audio, it has both research and pedagogical implications.
This paper presented a professionalized system-paper version of the B.E. project Automated Speaker Video Dubbing System. Based on the verified project report and runtime log, the work can be described accurately as a modular English-to-Hindi video-dubbing pipeline with demonstrated proof-of-concept execution. The project shows that source video handling, speech extraction, transcription, translation, Hindi speech generation, and final recombination can be connected into a working multimedia pipeline. The project should therefore be understood in the right way: not as a finished industrial dubbing product, but as a meaningful and well-scoped prototype. Its present strength lies in integration and execution feasibility. Its current limitations lie in naturalness, timing precision, reproducibility, and professional-quality delivery. Future work can focus on better context-aware translation, stronger Hindi speech synthesis, improved speaker preservation, more robust background-audio remixing, formal evaluation, cleaner deployment, and carefully validated lip-sync extensions.
A Modular English-to-Hindi Video Dubbing Pipeline: Design, Implementation, and Proof-of-Concept Evaluation · 2026 · DOIThis study recommends conducting further research on AD in the Arab world to expand the accessibility services provided by official TV channels and streaming platforms.
Aspects of Visual Content Covered in the Audio Description of Arabic Series: A Corpus-assisted Study · 2023 · DOIIt focuses on how functional Cantonese–English–Japanese multilinguals with partially overlapping language systems encode and gauge similarity of voluntary motion in their L1, which is a rarely studied language combination.
Multilingual learning and cognitive restructuring: The role of audiovisual media exposure in Cantonese–English–Japanese multilinguals’ motion event cognition · 2022 · DOITo address this gap in the literature, the authors of this paper surveyed classroom audiovisual (AV) support professionals from 49 ARL institutions and conducted seven follow-up interviews.
It examines the multimedia learning and multimedia language learning theories that underlie the MVA research, synthesizes the findings on MVA in the last decade, and identifies three underresearched areas on the subject.
The benefits of using video and captions for improving general L2 reading and listening comprehension have been well documented, however what is lacking is research that explores what contribution they may make to learning beyond just comprehension.
The Effect and the Influence of the Use of Video and Captions on Second Language Learning · 2008Are audio media more effective than visual media? Is a picture worth 1, 000 words? Research studies in this area are inconclusive and subject to numerous inter pretive difficulties; however, there is general agreement that the use of both media reinforces the transmission of a message (DeBoth and Dominowski 1978).
The probability limits at the 1 and 5 per cent levels of confidence there if Downloaded by [Harvard Library] at 10:49 06 October 2014 SEMANTIC DIFFERENTIAL FOR THEATRE CONCEPTS 7 remain to be determined for the Theatre Semantic Differential as a whole, for the various factors, and for the in- dividual scales.
In both motion pictures and television it would seem that relatively greater attention should be paid to the video. If these media of mass communication have an advantage peculiar to themselves, it is in the use of the picture. With the advantage shown in this study for the video, in spite of the fact that the film was biased in favor of the audio, it would seem especially important that greater emphasis be given to the video in planning and producing television proinstructional films or grams. 2. That the impact of the audio and video portions of the films does not at times seem to be additive is perhaps due to the fact that although the sound track and picture are shown simultaneously the content of each is not always closely related, the one to the other. Film and television producers should be sure that the picture and sound track are well integrated, so as to reinforce one another.
With the increase in the use of face masks as a health precaution in the post-pandemic era, the effects of wearing such masks on language learners’ speech perception, which depends highly on visual cues, remain uncertain.
Unmasking language learning: impact of wearing face masks on the listening comprehension and word recognition of EFL learners in Saudi Arabia · 2025 · DOIConsequently, the accuracy, reliability, and educational value of health-related videos remain uncertain, especially for complex procedures such as awake brain surgery involving language mapping.
YouTube As a Source of Information on Awake Brain Surgery and Language: Quality, Reliability, And Professional Involvement · 2025 · DOIWhereas several studies have explored the expression of emotions, little is known on how the visual and audio channels are combined during production of what we call the more controlled social affects, for example, “attitudinal” expressions.
Abstract Despite the widespread use of digital video and, increasingly, multimedia in listening instruction throughout second language programs, little is known about how learners attend to dynamic visual elements in comprehension.
The data, however, are inconclusive in establishing a relationship between the extent to which such a strategy is employed and performance in translation, as measured by amount of omitted material.
Most-cited papers in Subtitles and Audiovisual Media
- The Effect of Imagery and On‐Screen Text on Foreign Language Vocabulary Learning From Audiovisual Input · TESOL Quarterly · 2019 · 115 citations
- To live (code) or to not: A new method for coding in qualitative research · Qualitative Social Work · 2019 · 103 citations
- A state-of-the-art review of the modes and effectiveness of multimedia input for second and foreign language learning · Computer Assisted Language Learning · 2021 · 67 citations
- Performing Qualitative Content Analysis of Video Data in Social Sciences and Medicine: The Visual-Verbal Video Analysis Method · International Journal of Qualitative Methods · 2023 · 50 citations
- Investigating the Role of <scp>AI</scp> Tools in Enhancing Translation Skills, Emotional Experiences, and Motivation in <scp>L2</scp> Learning · European Journal of Education · 2024 · 48 citations
- Netflix Originals in Spain: Challenging diversity · European Journal of Communication · 2021 · 44 citations
- Synthetic versus human voices in audiobooks: The human emotional intimacy effect · New Media & Society · 2021 · 44 citations
- Understanding Digitized Second Language Videotext · Computer Assisted Language Learning · 2004 · 39 citations
- “Suit the Action to the Word, the Word to the Action”: An Unconventional Approach to Describing Shakespeare's <i>Hamlet</i> · Journal of Visual Impairment & Blindness · 2009 · 37 citations
- The effects of using an auto-subtitle system in educational videos to facilitate learning for secondary school students: learning comprehension, cognitive load, and satisfaction · Smart Learning Environments · 2023 · 35 citations
Most recent work
- The effects of audiovisual input on second language learning: <i>A meta-analysis</i> – ERRATUM · Studies in Second Language Acquisition · 2026
- A Modular English-to-Hindi Video Dubbing Pipeline: Design, Implementation, and Proof-of-Concept Evaluation · Zenodo (CERN European Organization for Nuclear Research) · 2026
- Audiovisual (a)synchrony · Studies in Second Language Acquisition · 2026
- Designing an Inclusive Prototype for Audience Feedback Collection Evaluation Tool in Cultural Contexts and Live Events: A Case Study at the Museo Tattile Statale Omero · Culture · 2026
- The impact of subtitle speed and number of lines on saccadic latency: An eye-tracking study ETRA025 · Proceedings of the ACM on Human-Computer Interaction · 2026
- Audible Sense: Turning Emotionally Adaptive Subtitles into Human-like Speech · International Journal for Research in Applied Science and Engineering Technology · 2026
- KÜNZLI, Alexander and KAINDL, Klaus. Handbuch Audiovisuelle Translation. Arbeitsmittel für Wissenschaft, Studium, Praxis. Berlin, Frank & Timme, 2024, 429 pp., ISBN 978-3-7329-8957-7 · Hikma · 2026
- MULTIMODAL APPROACH TO THE ORAL TRANSLATION OF VIDEO MATERIALS IN CONTEMPORARY MEDIA ENVIRONMENTS · Наукові інновації та передові технології · 2026
- ‘ <i>Deaf Is Only One of Us</i> ’ and Other Viewpoints in Historical Debates on TV and Film Captioning in Hong Kong · Sociology Lens · 2026
- Exploring the phrase-internal changes in articulation rate: the LARometer tool and its applications · AUC PHILOLOGICA · 2026
Find a gap in your own Subtitles and Audiovisual Media sub-topic
This page shows what the Subtitles and Audiovisual Media literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →