Arts and Humanities · Research topic

Open research questions in Subtitles and Audiovisual Media

29 unresolved questions extracted from the limitations and future-work sections of 896 Subtitles and Audiovisual Media papers in our library. Each links back to the study that raised it.

What the literature leaves open

  • However, free online corpora of video game dialogue in languages besides English and Japanese (such as the “FIGS” languages—French, Italian, German, and Spanish) are lacking.

    A multi-language video game dialogue corpus · 2026
  • For Learners: Because watching movies frequently with English subtitles is strongly linked to better TOEIC listening performance, students should transition away from Thai subtitles. Watching English-language films three to four times a week is recommended for steady improvement in listening skills. For Teachers: Instead of simply telling students to watch English movies, instructors should actively promote the use of English subtitles and create post-viewing activities, like dictation or comprehension quizzes, to solidify learning. Since Netflix is highly popular among students, teachers could also curate a list of shows on the platform that align with their students' skill levels. For Future Researchers: Upcoming studies should explore the motivations driving students' subtitle preferences—specifically, whether they use them out of habit or a deliberate desire to learn—as this context is crucial for interpreting data. Additionally, employing longitudinal methods and more diverse participant groups would help make the findings more universally applicable.

    Relationship between Behaviors in Watching English Movies and English Listening Skill · 2026 · DOI
  • The current study is subject to several limitations. First, it only evaluated vocabulary learning at the level of form recall. To gain a more comprehensive picture of how textual enhancement in captions affects incidental vocabulary learning, future research can use several tests to measure vocabulary development at different sensitivity levels (form recognition, meaning recall/recognition). Second, a transferappropriate processing concern arises from the present design. According to the transfer-appropriate processing framework (Morris et al., 1977), learning performance is influenced by the congruency between the conditions of learning and the conditions of testing. Since learners in the captioned groups had access to written target forms during treatment while the uncaptioned group did not, the written posttest format may have systematically advantaged the captioned groups due to learning-testing congruency, making it difficult to determine the extent to which group differences reflected vocabulary learning rather than test-format effects. Third, the initial letters of the target items were supplied as clues to prevent learners from writing alternative items. This differs from L2 users’ production of vocabulary items in real-life situations. Fourth, the study was conducted at a single site with English-major university students, limiting generalisability. Vitta et al. (2022) highlighted the importance of multisite designs in L2 vocabulary research. Fifth, English-major students represent a specialised learner profile with potentially different characteristics than non-major learners. Future multisite research with more diverse learner profiles is therefore warranted. Sixth, the sample size per condition was relatively small. However, it should be mentioned that previous research on the effect of mode of input on incidental vocabulary learning has used the same number of participants, at least in one of their experimental conditions (Feng & Webb, 2020; Majuddin et al., 2021). Seventh, word- and learner-related variables, such as frequency, working memory, and visual support (see Teng & Cui, 2025; Kaderoğlu, 2026), were not investigated. Finally, to obtain a more fine-grained picture of the effect of the treatment conditions, future studies are recommended to include a test-only control group.

    The Effect of Textual Enhancement on Incidental Learning of Single Words and Multi-word Items From L2 Captioned Viewing · 2026 · DOI
  • However, with no concurrent studies to cope up and update this recent convenience of the TTS, little is known about its scale and forms of impact on the population of the Arabic speaking EFL learners.

    Revisiting the Text to Speech Tool in EFL Learning in the Light of the Recent Rise in the Accessibility of the Technology · 2026 · DOI
  • This study had a number of limitations. It included only immediate and receptive learning posttests of form and form-meaning mapping, which limits any claims we can make regarding durability of learning, and did not include a test of word production. Additionally, the type of mapping selected (new L2 form to existing L1 concept) was specific within the broad possibilities for form-meaning mappings in L2 vocabulary learning (see Kida & Barcroft, 2018). As such, generalizable claims regarding learning are limited to L2 form-form and form-meaning mapping. One anonymous reviewer noted that we included both general and specific distractor types on the meaning recognition outcome test, tailored to the meaning associated with that pseudoword item, which could have impacted item responses; 76% of meaning recognition items (19/25) included only specific distractors, 12% (3/25) included only general distractors, and 12% (3/25) had one of each. Descriptive scores indicated notable differences in accuracy between items with specific (58% accurate), generic (67% accurate), or mixed (38% accurate) distractors, so we re-ran the main outcome analysis for the meaning recognition measure with distractor type (specific, generic, mixed) as a categorical predictor. In this analysis, distractor type was not a significant predictor of meaning recognition when other variables were included (b = (cid:1).41, SE =.27, p =.127), and did not change other results of the model, so the model without distractor type was retained as the best-fitting model. However, future studies utilizing similar outcome presentation modalities and examining learning of only specific or general items can better control for this potential source of variability in learning. The study was highly controlled, with pseudoword targets rather than real words, while the audio rate was standardized rather than modified to follow reading speed more closely at the individual level. Reading ahead was defined dichotomously yes/no in the analytic models, so it was not possible to examine the relative effects of lagging behind the audio when reading. Since our RQs focused only on the potential benefits of reading ahead, we did not explore differential impacts of reading precisely synchronously or lagging behind the audio; future studies are warranted which do so, thereby increasing precision and allowing for more specific theoretical predictions for instances when the eye and ear are synchronous, and when the eye lags behind the ear during RWL. The L1s among participants were diverse (see Table 2), which did not allow for an examination of the effects by L1, while recent findings from L2 English reading report evidence that L1-L2 differences may result in variable reading behavior, demonstrated by eye movements (Kuperman et al., 2025). This would be a fruitful area for future research. As the present study provides initial evidence that reading ahead of the audio during RWL facilitates faster word form recognition across instances and mapping form to meaning during contextualized learning of new L2 vocabulary, compared with reading synchronously or behind the audio, it has both research and pedagogical implications.

    Audiovisual (a)synchrony · 2026 · DOI
  • This paper presented a professionalized system-paper version of the B.E. project Automated Speaker Video Dubbing System. Based on the verified project report and runtime log, the work can be described accurately as a modular English-to-Hindi video-dubbing pipeline with demonstrated proof-of-concept execution. The project shows that source video handling, speech extraction, transcription, translation, Hindi speech generation, and final recombination can be connected into a working multimedia pipeline. The project should therefore be understood in the right way: not as a finished industrial dubbing product, but as a meaningful and well-scoped prototype. Its present strength lies in integration and execution feasibility. Its current limitations lie in naturalness, timing precision, reproducibility, and professional-quality delivery. Future work can focus on better context-aware translation, stronger Hindi speech synthesis, improved speaker preservation, more robust background-audio remixing, formal evaluation, cleaner deployment, and carefully validated lip-sync extensions.

    A Modular English-to-Hindi Video Dubbing Pipeline: Design, Implementation, and Proof-of-Concept Evaluation · 2026 · DOI
  • This study recommends conducting further research on AD in the Arab world to expand the accessibility services provided by official TV channels and streaming platforms.

    Aspects of Visual Content Covered in the Audio Description of Arabic Series: A Corpus-assisted Study · 2023 · DOI
  • It focuses on how functional Cantonese–English–Japanese multilinguals with partially overlapping language systems encode and gauge similarity of voluntary motion in their L1, which is a rarely studied language combination.

    Multilingual learning and cognitive restructuring: The role of audiovisual media exposure in Cantonese–English–Japanese multilinguals’ motion event cognition · 2022 · DOI
  • To address this gap in the literature, the authors of this paper surveyed classroom audiovisual (AV) support professionals from 49 ARL institutions and conducted seven follow-up interviews.

    The Future of Video Playback Capability in College and University Classrooms · 2017 · DOI
  • It examines the multimedia learning and multimedia language learning theories that underlie the MVA research, synthesizes the findings on MVA in the last decade, and identifies three underresearched areas on the subject.

    Using Multimedia Vocabulary Annotations in L2 Reading and Listening Activities · 2010 · DOI
  • The benefits of using video and captions for improving general L2 reading and listening comprehension have been well documented, however what is lacking is research that explores what contribution they may make to learning beyond just comprehension.

    The Effect and the Influence of the Use of Video and Captions on Second Language Learning · 2008
  • Are audio media more effective than visual media? Is a picture worth 1, 000 words? Research studies in this area are inconclusive and subject to numerous inter pretive difficulties; however, there is general agreement that the use of both media reinforces the transmission of a message (DeBoth and Dominowski 1978).

    The Effectiveness of Audio-Visual Presentations · 1981 · DOI
  • The probability limits at the 1 and 5 per cent levels of confidence there if Downloaded by [Harvard Library] at 10:49 06 October 2014 SEMANTIC DIFFERENTIAL FOR THEATRE CONCEPTS 7 remain to be determined for the Theatre Semantic Differential as a whole, for the various factors, and for the in- dividual scales.

    A semantic differential for theatre concepts · 1961 · DOI
  • In both motion pictures and television it would seem that relatively greater attention should be paid to the video. If these media of mass communication have an advantage peculiar to themselves, it is in the use of the picture. With the advantage shown in this study for the video, in spite of the fact that the film was biased in favor of the audio, it would seem especially important that greater emphasis be given to the video in planning and producing television proinstructional films or grams. 2. That the impact of the audio and video portions of the films does not at times seem to be additive is perhaps due to the fact that although the sound track and picture are shown simultaneously the content of each is not always closely related, the one to the other. Film and television producers should be sure that the picture and sound track are well integrated, so as to reinforce one another.

    The relative contribution to learning of video and audio elements in films · 1951 · DOI
  • With the increase in the use of face masks as a health precaution in the post-pandemic era, the effects of wearing such masks on language learners’ speech perception, which depends highly on visual cues, remain uncertain.

    Unmasking language learning: impact of wearing face masks on the listening comprehension and word recognition of EFL learners in Saudi Arabia · 2025 · DOI
  • Consequently, the accuracy, reliability, and educational value of health-related videos remain uncertain, especially for complex procedures such as awake brain surgery involving language mapping.

    YouTube As a Source of Information on Awake Brain Surgery and Language: Quality, Reliability, And Professional Involvement · 2025 · DOI
  • Whereas several studies have explored the expression of emotions, little is known on how the visual and audio channels are combined during production of what we call the more controlled social affects, for example, “attitudinal” expressions.

    Multimodal Indices to Japanese and French Prosodically Expressed Social Affects · 2009 · DOI
  • Abstract Despite the widespread use of digital video and, increasingly, multimedia in listening instruction throughout second language programs, little is known about how learners attend to dynamic visual elements in comprehension.

    Understanding Digitized Second Language Videotext · 2004 · DOI
  • The data, however, are inconclusive in establishing a relationship between the extent to which such a strategy is employed and performance in translation, as measured by amount of omitted material.

    Simultaneous Interpretation: Temporal and Quantitative Data · 1973 · DOI

Most-cited papers in Subtitles and Audiovisual Media

Most recent work

Find a gap in your own Subtitles and Audiovisual Media sub-topic

This page shows what the Subtitles and Audiovisual Media literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.

Open the Research Gap Finder →

Related topics in Arts and Humanities

29 open questions have been extracted from the limitations and future-work passages of 896 Subtitles and Audiovisual Media papers in our library. Each one below links back to the study that raised it, so you can read the original claim in context.

Tools for your next paper

Compare the categoryHonest roundups of the AI research tools, ours listed alongside the alternatives.

Command palette

Jump anywhere, run any action.