Open research questions in Music and Audio Processing
52 unresolved questions extracted from the limitations and future-work sections of 285 Music and Audio Processing papers in our library. Each links back to the study that raised it.
What the literature leaves open
The cold start challenge is a significant problem in existing music recommendation systems. Traditional methods neglect the inherent attributes of songs, focusing primarily on user behaviour. The proposed system aims to address these challenges by combining Fuzzy C-means clustering and cosine similarity.
Future work may involve incorporating recurrent neural networks (RNNs) for better temporal feature analysis. The system can be extended to support multilingual audio analysis and real-time streaming voice authentication. Additional biometric features such as speaker emotion, speech rhythm, and voice behavior analysis can be incorporated to strengthen authentication capability.
The creation of high-quality deepfake audio has become more accessible with the advent of generative models. There is a need for effective detection mechanisms to combat deepfake audio challenges.
The difficulty of soundchecking in small venues - The need for a listener with music experience - The challenge of scaling the framework to other instruments and control parameters
Generation latency and unpredictable outputs remain challenges for neural text-to-audio systems. The tweakability problem for neural audio systems remains a challenge. Sharp stylistic pivots and timbral specificity are difficult to achieve.
The complexity of Indian classical music - The limited availability of labeled datasets - The need for accurate and efficient methods for raga recognition
RagaDetNet: deep learning based Indian classical raga recognition using improved Archimedes optimization algorithm for acoustic feature selection · 2026 · DOIApplying the proposed approach to other sequential data such as speech or text. Exploring the use of other temporal-style descriptors such as wavelet transforms or Fourier analysis.
Explicit encoding of temporal dynamics for interpretable melody similarity modeling: insight into machine learning models · 2026 · DOICurrent methods are mostly based on implicit encoding of temporal dynamics. There is a need for a model-agnostic feature representation approach that enables controlled comparison and interpretability.
Explicit encoding of temporal dynamics for interpretable melody similarity modeling: insight into machine learning models · 2026 · DOIWhile semantic priors are widely exploited to enhance linguistic intelligibility, the integration of explicit acoustic priors remains underexplored, limiting synthesis fidelity in frequency-sensitive domains.
MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation · 2026This exploratory research provides initial theoretical support for further investigation of music listening engagement utilising an SDT framework.
Using Generative Artificial Intelligence to Identify Themes of Basic Psychological Needs in Popular Song Lyrics · 2026 · DOIThe mechanisms underlying how music listening influences well-being remain unclear, especially in terms of broader motivational processes.
Using Generative Artificial Intelligence to Identify Themes of Basic Psychological Needs in Popular Song Lyrics · 2026 · DOIWhile the field has advanced significantly on monophonic and piano-form scores, multi-part score transcription remains underexplored, largely due to the absence of a suitable dataset.
A Dataset and Benchmark for Optical Music Recognition of String Quartet Scores · 2026Data sources in a dataset cannot be mixed and matched at this time; for example, a dataset built from a GitHub repository cannot include files uploaded from a local system, limiting data integration flexibility.
An Intelligent Music Recommendation System Using Machine Learning and User Preference Analysis · 2026 · DOIThe "Opposite Song Generation Process" is introduced to address the challenges associated with music exploration, such as breaking free from recommendation loops and encouraging users to discover songs outside their comfort zones. By deliberately selecting songs with contrasting characteristics, the model enriches the user experience, promoting better exploration of musical tastes. In summary, the opposite song selection process adds a layer of innovation to the music recommendation model, ensuring that users not only receive recommendations aligned with their preferences but also have the opportunity to explore and appreciate a broader spectrum of musical expressions. The implementation realizes the proposed model, using essential libraries such as pandas, numpy, fcmeans, and cosine_similarity. The Flask framework enables the creation of a user-centric web application, seamlessly integrating recommendation and opposite song generation functionalities. The code encapsulates functions tailored for user request handling, data processing, clustering execution, similarity score computation, and the provisioning of recommended and opposite songs in JSON format. Figure 3: Visualisation of Other Side of the Coin recommendation in the embedding space, where songs most dissimilar to the given song are chosen from all other clusters The red dot in figure 3 symbolises the input song, and the black dot indicates the songs recommended to the user. In this situation, when the user inputs a song (depicted by the red dot), the clusters, excluding the one to which the song belongs, are sorted using cosine similarity relative to the input song. This approach aims to provide the user with song recommendations that offer diversity by suggesting tracks that differ from the user's preferences.
Evaluating the effectiveness of Visual Transfer Learning on other datasets and domains. Comparing the performance of Visual Transfer Learning to other approaches, such as traditional machine learning methods.
The lack of effective methods for Environmental Sound Classification in data-constrained scenarios. The need for more accurate and efficient ESC systems.
To improve the accuracy of the emotion detection algorithms. To incorporate more modalities, such as physiological signals, into the system. To develop a more personalized and adaptive music recommendation system.
The conventional technique of recommending music does not take into account the current emotional status of the user. Human emotions are intricate and dynamic, and algorithms relying on one data source are insufficient to represent humans' emotional states.
Existing approaches have limitations such as weak cross-genre generalization. Insufficient modeling of long-range temporal dependencies. Inadequate capture of hierarchical emotional structures.
A deep learning framework for emotion recognition in music using multimodal data fusion · 2026 · DOITo enable real-time deployment in user-facing applications such as streaming music platforms or emotion-aware recommendation systems, future work could explore lightweight variants of our architecture using model pruning, quantization, or knowledge distillation. Techniques such as domain adaptation, few-shot learning, or semi-supervised multimodal training could be explored in this context.
A deep learning framework for emotion recognition in music using multimodal data fusion · 2026 · DOIPersistent gaps remain in cultural inclusivity, interpretability, and ethical governance. The study excludes pure music theory research, art criticism without technical method support, and nonacademic format content.
Artificial Intelligence in Music: A Bibliometric and Systematic Review of Creation, Performance, and Education · 2026 · DOIFuture models are expected to achieve heightened emotional and contextual awareness, powering hyper-personalized systems that adapt to users' physiological and behavioral cues. Emerging trends point to immersive, interactive, and context-responsive systems. The study suggests a growing research focus on multimodal perception systems, where AI music interacts dynamically with environmental and visual inputs.
Artificial Intelligence in Music: A Bibliometric and Systematic Review of Creation, Performance, and Education · 2026 · DOIFuture research could employ silhouette analysis or elbow-method diagnostics to discover the unconstrained optimal clustering structure of the catalogue.
Quantitative Evaluation of Stylistic Consistency and Acoustic Evolution: A Case Study of Jay Chou's Career from 2000 to 2026 · 2026 · DOILong-term acoustic analysis of individual artists remains relatively limited in Mandopop studies. Longitudinal quantitative studies of individual artists in Mandopop remain relatively limited.
Quantitative Evaluation of Stylistic Consistency and Acoustic Evolution: A Case Study of Jay Chou's Career from 2000 to 2026 · 2026 · DOIExisting approaches face significant limitations, including single-modality methods and overlooking interactive relationships. A single modality may not be sufficient to capture the diversity of a music tag repository.
A multimodal graph-based music auto-tagging framework: integrating social and content intelligence · 2026 · DOI
Most-cited papers in Music and Audio Processing
- AudioLDM 2: Learning Holistic Audio Generation With Self-Supervised Pretraining · IEEE/ACM Transactions on Audio Speech and Language Processing · 2024 · 171 citations
- Transformer and Graph Convolution-Based Unsupervised Detection of Machine Anomalous Sound Under Domain Shifts · IEEE Transactions on Emerging Topics in Computational Intelligence · 2024 · 92 citations
- Exploring Deep Learning Methods for Audio Speech Emotion Detection: An Ensemble MFCCs, CNNs and LSTM · Applied Mathematics & Information Sciences · 2024 · 65 citations
- More of the Same – On Spotify Radio · Culture Unbound Journal of Current Cultural Research · 2017 · 15 citations
- AI-Powered Choreography Using a Multilayer Perceptron Model for Music-Driven Dance Generation · Informatica · 2025 · 8 citations
- INTELLIGENT MUSIC APPLICATIONS: INNOVATIVE SOLUTIONS FOR MUSICIANS AND LISTENERS · Uluslararası Anadolu Sosyal Bilimler Dergisi · 2023 · 7 citations
- Wind Sounds Classification Using Different Audio Feature Extraction Techniques · Informatica · 2022 · 7 citations
- Evaluating Similarity of Spectrogram-like Images of DC Motor Sounds by Pearson Correlation Coefficient · Elektronika ir Elektrotechnika · 2022 · 7 citations
- Solfeggio Teaching Method Based on MIDI Technology in the Background of Digital Music Teaching · International Journal of Web-Based Learning and Teaching Technologies · 2023 · 7 citations
- Optimization of Piano Performance Teaching Mode Using Network Big Data Analysis Technology · International Journal of Information and Communication Technology Education · 2024 · 7 citations
Most recent work
- MSA-TCN: Robust Urban Suspicious Sound Detection Using Multi-Scale Temporal Convolutions and Dual Attention · Journal of Future Artificial Intelligence and Technologies · 2026
- Artificial Intelligence in Music: A Bibliometric and Systematic Review of Creation, Performance, and Education · Journal of Artificial Intelligence and Soft Computing Research · 2026
- An Intelligent Music Recommendation System Using Machine Learning and User Preference Analysis · International Journal for Research in Applied Science and Engineering Technology · 2026
- ScreamAlert: A TinyML-Powered Wearable for Instant Acoustic Emergency Detection · INTERNATIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT · 2026
- Comparative Analysis of SVM, k-NN and Logistic Regression Methods in Classifying Turkish Music Genres · Journal of Innovative Science and Engineering (JISE) · 2026
- A Comprehensive Review of Audio Segmentation and Classification: from Rule-Based Methods to YOHO and Deep Learning Approaches · Archives of Computational Methods in Engineering · 2026
- Unveiling Digital Drugs Through Artificial Intelligence: An Audio Feature Analysis Approach · International Journal of Intelligent Engineering and Systems · 2026
- Characterizing continual learning scenarios and strategies for audio analysis · Journal on Audio, Speech, and Music Processing · 2026
- Methods for Pitch Analysis in Contemporary Popular Music: Multiple Pitches From Harmonic Tones in Vitalic’s Music · Journal of the Audio Engineering Society · 2026
- DGSNA: Dynamic Generative Scene-Based Noise Addition Method · Computation · 2026
Find a gap in your own Music and Audio Processing sub-topic
This page shows what the Music and Audio Processing literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →