Open research questions in Imbalanced Data Classification Techniques
59 unresolved questions extracted from the limitations and future-work sections of 327 Imbalanced Data Classification Techniques papers in our library. Each links back to the study that raised it.
What the literature leaves open
Future work will explore evaluation using synthetic tax micro- simulation environments and privacy-controlled regulatory sandboxes to better approxi- mate real taxation ecosystems. A limitation of the current study is the absence of large-scale confidential tax au- dit datasets, which restricts evaluation of behavioural fraud patterns and policy-driven audit triggers.
Privacy preserving and auditable tax fraud detection using zero knowledge transformer inference · 2026 · DOITo close the gap between current research and real banking use, future work should prioritize these directions—ordered by frequency and severity of challenges in Section 4: 1) Privacy-preserving FL and multi-national data sharing (top priority)—directly tackle data quality, privacy, and generalizability issues seen in more than 30% of the papers, while handling different international rules. 2) Lightweight models for edge deployment (e.g., TinyML)— address the most common complaint: computational demands and scalability problems in real-time processing. FinTech and Sustainable Innovation Vol. 00 Iss.
A Systematic Review of AI-Driven Banking Fraud Detection: Advances, Challenges, and Deployment-Ready Solutions (2024–2025) · 2026 · DOIThis study created a DSS to identify health insurance fraud. It makes use of a hybrid framework that combines GSVM with genetic algorithms.
Decision Support System (DSS) for Fraud Detection in Health Insurance Claims using Genetic Support Vector Machines (GSVMS) · 2026 · DOIFuture applications of similar systems should incorporate a real payment gateway instead of simulated transaction data from the beginning of the project. Training datasets need to be continually increased and updated to ensure that the performance of the model remains accurate as fraud and review language patterns change over time. Developers need to make sure that their pipelines are continuous for retraining, not just one-time deployments. Any company that plans to implement a similar design should test extensively and in a staging environment before deployment, especially the fraud detection element, which seems to have operational costs both from false positives and false negatives. REFERENCES Banu, R., Ashok, A., Dwivedi, V. K., Reddy, K. A., Thulasimani, T., & Nishant, N. (2024). An innovative method for fraud detection in e-commerce using DCNN-multiclass SVM model. In 2024 International Conference on Intelligent Algorithms for Computational Intelligence Systems (IACIS) (pp. 1-6). IEEE. https://doi.org/10.1109/IACIS61494.2024.10721774 Coherent Market Insights. (2023). Fashion e-commerce market analysis and growth projections. Retrieved from https:// www.coherentmarketinsights.com Kaggle. (2023a). Credit card fraud detection dataset. Retrieved https://www.kaggle.com/datasets/mlg-ulb/ from creditcardfraud Kaggle. (2023b). Women's e-commerce clothing reviews dataset. Retrieved from https://www.kaggle.com/datasets/ nicapotato/womens-ecommerce-clothing-reviews Kumar, S., Gunjan, V. K., Ansari, M. D., & Pathak, R. (2022). Credit card fraud detection using support vector machine. In Proceedings of the 2nd International Conference on Recent Trends in Machine Learning, IoT, Smart Cities and Applications: ICMISC 2021 (pp. 27-37). Springer, Singapore. https://doi.org/10.1007/978-981-16-6407-6_3 Mutemi, A., & Bacao, F. (2024). E-commerce fraud detection based on machine learning techniques: Systematic literature review. Big Data Mining and Analytics, 7(2), 419-444. https:// doi.org/10.26599/BDMA.2023.9020024 Sharma, H. D., & Goyal, P. (2024). Interpretable aspect based sentiment classification of online educational reviews using SVM model and explainable LIME-AI model. International Journal of Information Technology, 16, 4567-4578. https://doi. org/10.1007/s41870-024-02125-w Shopify. (2024). E-commerce statistics and trends: Social commerce revenue data. Retrieved from https://www. shopify.com/research Statista. (2023). Global fashion e-commerce market size and forecasts 2024-2030. Retrieved from https://www.statista.com Tabany, M., & Gueffal, M. (2024). Sentiment analysis and fake Amazon reviews classification using SVM supervised machine learning model. Journal of Advances in Information Technology, 15(1), 49-58.
Integrating Support Vector Machine Classifiers for Real-Time Sentiment Analysis and Fraud Detection in A Fashion E-Commerce Platform · 2026 · DOIThere are a number of extensions that would be useful in the system. Linking the fraud detection module to a real paying gateway like Paystack or Flutterwave would replace the simulated feature vector with real financial transaction data, resulting in a much more reliable fraud classifier in production. The sentiment model could be used in more languages and with more context dependent expressions like sarcasm and that would make it a more useful model in markets outside of the English speaking one. A product recommendation system based on browsing and purchase history would help support the existing review and fraud features. Lastly, the application would be deployed in a real production environment, with the appropriate security hardening, for a thorough performance assessment in a real environment with load. RECOMMENDATIONS Future applications of similar systems should incorporate a real payment gateway instead of simulated transaction data from the beginning of the project. Training datasets need to be continually increased and updated to ensure that the performance of the model remains accurate as fraud and review language patterns change over time. Developers need to make sure that their pipelines are continuous for retraining, not just one-time deployments. Any company that plans to implement a similar design should test extensively and in a staging environment before deployment, especially the fraud detection element, which seems to have operational costs both from false positives and false negatives. REFERENCES Banu, R., Ashok, A., Dwivedi, V. K., Reddy, K. A., Thulasimani, T., & Nishant, N. (2024). An innovative method for fraud detection in e-commerce using DCNN-multiclass SVM model. In 2024 International Conference on Intelligent Algorithms for Computational Intelligence Systems (IACIS) (pp. 1-6). IEEE. https://doi.org/10.1109/IACIS61494.2024.10721774 Coherent Market Insights. (2023). Fashion e-commerce market analysis and growth projections. Retrieved from https:// www.coherentmarketinsights.com Kaggle. (2023a). Credit card fraud detection dataset. Retrieved https://www.kaggle.com/datasets/mlg-ulb/ from creditcardfraud Kaggle. (2023b). Women's e-commerce clothing reviews dataset. Retrieved from https://www.kaggle.com/datasets/ nicapotato/womens-ecommerce-clothing-reviews Kumar, S., Gunjan, V. K., Ansari, M. D., & Pathak, R. (2022). Credit card fraud detection using support vector machine. In Proceedings of the 2nd International Conference on Recent Trends in Machine Learning, IoT, Smart Cities and Applications: ICMISC 2021 (pp. 27-37). Springer, Singapore. https://doi.org/10.1007/978-981-16-6407-6_3 Mutemi, A., & Bacao, F. (2024).
Integrating Support Vector Machine Classifiers for Real-Time Sentiment Analysis and Fraud Detection in A Fashion E-Commerce Platform · 2026 · DOIThe 28-day window offered a useful balance between early intervention and predictive gain, but vari- ation across group-wise splits indicates that deployment decisions should be validated under the course structures and resource constraints of the target institution.
Calibrated early-warning models with fairness auditing and selective prediction for course withdrawal risk: Evidence from OULAD · 2026 · DOIDespite this, AML detection systems remain largely underexplored from a fairness perspective, even though deeper analytical methods based on counterfactuals are now available.
Counterfactual Methods for Detecting Unfairness in Anti-Money Laundering Algorithms · 2026CONCLUSION AND FUTURE SCOPE This review has argued that predictive performance alone is insufficient for high-stakes financial fraud detection: a deployable system must also be feature interpretable, auditable, efficient, and robust to changing transaction behaviour.
Explainable Credit Card Fraud Detection Using LightGBM And SHAP-Guided Feature Selection: A Review · 2026 · DOIFuture research will address the development of production-ready models and consider the implications of ML-supported fraud detection in the domain of insurance. These principles facilitate the scalable deployment of models tailored for the constant battle against ever-evolving fraudulent techniques, but scalability and operationalization remain to be proven.
Future research will explore streaming fraud detection with online XGBoost, graph neural networks for transactional network modelling, federated learning for privacy-preserving cross-institutional fraud detection, and adaptive retraining strategies to address concept drift in production deployments.
In this work, an effective online payment fraud detection system was developed using machine learning techniques. The combination of CatBoost and XGBoost models through an ensemble approach resulted in improved predictive performance. PCA was used for dimensionality reduction, and SMOTE was applied to address class imbalance, which significantly enhanced the model’s ability to detect fraudulent transactions. The system achieved high accuracy, precision, recall, and AUC scores, demonstrating its effectiveness in realworld scenarios. Additionally, the deployment of the model using Streamlit provides a practical interface for real-time fraud detection. Future work can focus on integrating deep learning approaches such as neural networks and graph-based models to capture more complex transaction patterns. Furthermore, real-time streaming data and largescale deployment can be explored to improve scalability and adaptability in dynamic financial environments. REFERENCES 1. M. Habibpour, H. Gharoun, M. Mehdipour, A. Tajally, H. Asgharnezhad, A. Shamsi, A. Khosravi, M. Shafie-Khah, S. Nahavandi, and J. P. S. Catalao, ''Uncertainty-aware Online payment fraud detection using deep learning 2021; arXiv:2107.13508. 2. A. Cherif, A. Badhib, H. Ammar, S. Alshehri, M. Kalkatawi, and A. Imine. "Online payment fraud detection in the era of disruptive technologies: A systematic review." J. King Saud Univ. Computer and Information Science, vol. 35, no. 1, pp. 145-174, Jan. 2023, doi:10.1016/j.jksuci.2022.11.008. 3. T. K. Dang, T. C. Tran, L. M. Tuan, and M. V. Tiep. "Machine learning based on resampling approaches and deep reinforcement learning for Online payment fraud detection systems." Appl. Sci., vol. 11, no. 21, p. 10004, Oct. 2021; doi: 10.3390/app112110004. Page 927 www.rsisinternational.org INTERNATIONAL JOURNAL OF LATEST TECHNOLOGY IN ENGINEERING, MANAGEMENT & APPLIED SCIENCE (IJLTEMAS) ISSN 2278-2540 | DOI: 10.51583/IJLTEMAS | Volume XV, Issue IV, April 2026 4. Chaquet-Ulldemolins et al., ''On the black-box problem for fraud detection using machine learning (I): Linear models and informative feature selection,'' Applied Sciences, vol. 12, no. 7, p. 3328, March 2022, doi: 10.3390/app12073328. 5. E. F. Malik, K. W. Khaw, B. Belaton, W. P. Wong, and X. Chew. "Online payment fraud detection using a new hybrid machine learning architecture." Mathematics, vol. 10, no. 9, p. 1480, April 2022; doi: 10.3390/math10091480. 6. I. Benchaji, S. Douzi, B. El Ouahidi, and J. Jaafari, "Enhanced Online payment fraud detection using attention mechanism and LSTM deep model," J. Big Data, vol. 8, no. 1, p. 151, December 2021; doi: 7. 10.1186/s40537-021-00541-8. 8. E. Esenogho, I. D. Mienye, T. G. Swart, K. Aruleba, and G. Obaido.
Data privacy and security enhancements such as encryption and access control mechanisms for the Kafka streaming pipeline and result streams have not been specified in detail; compliance validation with regulatory standards for transaction data remains unaddressed.
Integration of blockchain technology for improving transparency and tamper-resistance of transaction records in the Kafka-based fraud detection system has been suggested but no technical specifications, consensus mechanisms, or performance trade-offs have been defined or evaluated.
The system's scalability using Kubernetes and Docker containerization for managing multiple Kafka brokers and database instances has been mentioned conceptually but lacks empirical validation; performance benchmarks under varying transaction volume loads and latency constraints are absent.
Automated response capabilities such as account freezing and additional authentication triggers have been proposed but not integrated or tested within the real-time Kafka-based fraud detection architecture; the effectiveness of these proactive mechanisms in reducing actual fraud losses remains unvalidated.
The system currently processes game log data; integration with multi-source data including customer device information, geolocation data, and social network analysis for contextual fraud detection has been proposed but not experimentally validated for detecting subtle anomalies undetectable from transaction data alone.
Online learning and incremental learning techniques for continuously updating ML models as new transaction data arrives in the Kafka stream have not been implemented or benchmarked; the performance impact of continuous model adaptation versus periodic retraining on emerging fraud tactics remains unquantified.
Label confirmation delays in returns abuse detection (purchase-return linkage scenarios) have not been quantified or modeled as a function of return window duration, preventing optimization of supervised tabular classifiers when ground truth arrives weeks or months after the triggering event.
A Systematic Review of Machine Learning Approaches For AI-Driven Fraud Detection in Loyalty Programs · 2026 · DOIThe interaction between behavioral risk models trained on mobile-first session telemetry and account takeover detection after valid login has not been empirically validated across different redemption-lag windows, limiting understanding of optimal timing for behavioral anomaly scoring in the pre-redemption window.
A Systematic Review of Machine Learning Approaches For AI-Driven Fraud Detection in Loyalty Programs · 2026 · DOINo quantitative framework exists for choosing between interpretable tabular models and graph-based approaches based on program size and fraud taxonomy breadth, preventing loyalty operations with intermediate-scale datasets from objectively deciding when graph learning complexity is justified versus when simpler tabular models suffice.
A Systematic Review of Machine Learning Approaches For AI-Driven Fraud Detection in Loyalty Programs · 2026 · DOIThe sparse positive label problem in insider-assisted manipulation detection has not been addressed through specific semi-supervised or self-supervised learning techniques tailored to rare admin-action anomalies, leaving a gap in how to bootstrap models when labeled privileged-behavior cases are severely limited.
A Systematic Review of Machine Learning Approaches For AI-Driven Fraud Detection in Loyalty Programs · 2026 · DOIGraph-based models for mileage pooling and partner-transfer abuse require an integrated graph-ready data architecture, but the paper does not specify which graph construction methods, node/edge feature representations, or graph neural network architectures perform best when legacy loyalty systems lack native relational data infrastructure.
A Systematic Review of Machine Learning Approaches For AI-Driven Fraud Detection in Loyalty Programs · 2026 · DOICampaign-specific false-positive costs and their interaction with model drift in hybrid pipelines for promotional abuse detection have not been measured across different loyalty program rule-change frequencies, limiting understanding of when rolling retraining adequately addresses concept drift.
A Systematic Review of Machine Learning Approaches For AI-Driven Fraud Detection in Loyalty Programs · 2026 · DOIThreshold trade-offs between false-positive rates and reviewer capacity have not been estimated for specific loyalty fraud scenarios, leaving practitioners without quantitative guidance on how to calibrate detection thresholds for account takeover, fake account farming, and promotional abuse given finite analyst review resources.
A Systematic Review of Machine Learning Approaches For AI-Driven Fraud Detection in Loyalty Programs · 2026 · DOIThe paper does not empirically rank machine learning architectures for loyalty fraud detection on a shared loyalty dataset, preventing direct performance comparison across supervised tabular models, anomaly detection, graph neural networks, and behavioral risk models for different fraud scenarios.
A Systematic Review of Machine Learning Approaches For AI-Driven Fraud Detection in Loyalty Programs · 2026 · DOI
Most-cited papers in Imbalanced Data Classification Techniques
- A survey on imbalanced learning: latest research, applications and future directions · Artificial Intelligence Review · 2024 · 409 citations
- Fraud detection: A systematic literature review of graph-based anomaly detection approaches · Decision Support Systems · 2020 · 356 citations
- Double machine learning with gradient boosting and its application to the Big N audit quality effect · Journal of Econometrics · 2020 · 217 citations
- The receiver operating characteristic curve accurately assesses imbalanced datasets · Patterns · 2024 · 179 citations
- Enhancing Credit Card Fraud Detection: An Ensemble Machine Learning Approach · Big Data and Cognitive Computing · 2024 · 178 citations
- A Novel Data Augmentation Method Based on Denoising Diffusion Probabilistic Model for Fault Diagnosis Under Imbalanced Data · IEEE Transactions on Industrial Informatics · 2024 · 169 citations
- Data oversampling and imbalanced datasets: an investigation of performance for machine learning and feature engineering · Journal Of Big Data · 2024 · 162 citations
- Data engineering for fraud detection · Decision Support Systems · 2021 · 144 citations
- Transparency and Privacy: The Role of Explainable AI and Federated Learning in Financial Fraud Detection · IEEE Access · 2024 · 136 citations
- Generative models improve fairness of medical classifiers under distribution shifts · Nature Medicine · 2024 · 117 citations
Most recent work
- Early risk stratification for carbapenem resistance among Pseudomonas aeruginosa infected patients using a clinico-laboratory machine-learning model based on routine complete blood count parameters · Frontiers in Cellular and Infection Microbiology · 2026
- Hybrid Deep Learning Framework with Cat Swarm Optimization for Cloud-Based Financial Fraud Detection · Mathematics · 2026
- Intelligent Credit Card Fraud Detection · Zenodo (CERN European Organization for Nuclear Research) · 2026
- Comparison Of Methods For Handling Imbalanced Datasets In Improving Classification Algorithm Performance · CAUCHY Jurnal Matematika Murni dan Aplikasi · 2026
- Online Payments Fraud Detection Using Machine Learning · INTERNATIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT · 2026
- Comparative Study of Machine Learning Algorithms for Fraud Detection in Blockchain · International Journal of Advanced Research in Science Communication and Technology · 2026
- DEFENSE-BANK: Sea Horse Optimized Deep Learning Framework for Fraud Detection in the Banking Sector · International Journal of Computational Intelligence Systems · 2026
- FinVision: AI-Driven Financial Fraud Detection · INTERNATIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT · 2026
- Real-Time Bank Transaction Fraud Detection Using Kafka and Machine Learning · International Journal of Creative and Open Research in Engineering and Management · 2026
- A Systematic Review of Machine Learning Approaches For AI-Driven Fraud Detection in Loyalty Programs · International Journal of Intelligent Data and Machine Learning · 2026
Find a gap in your own Imbalanced Data Classification Techniques sub-topic
This page shows what the Imbalanced Data Classification Techniques literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →