Open research questions in Adversarial Robustness in Machine Learning
281 unresolved questions extracted from the limitations and future-work sections of 928 Adversarial Robustness in Machine Learning papers in our library. Each links back to the study that raised it.
What the literature leaves open
While this work demonstrates the effectiveness of rewriting methods and, particularly, OBBR across diverse models, attack types, and evaluation settings, several important directions remain. Firstly, investigating domain-specific benign corpora which more closely align with safety-critical instruction-following data could enhance OBBR’s ability to filter subtle malicious patterns that do not rely on explicit triggers. Secondly, incorporating OBBR into other safety post-training phases—such as Safe RLHF (Dai et al., 2024) or SafeDPO (Anonymous, 2026)—could provide end-to-end poisoning protection throughout the model development lifecycle. Finally, exploring model-internal rewriting mechanisms may lead to new safety-enhancing architectures which further improve safety in the face of BAs and PIAs.
Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks · 20264 Ablation Study for Regularization Parameter λ In the black-box setting, the injection rate ρ directly controls how much poisoned data is mixed into the clean training set, while in the white-box setting the regularization parameter λ plays an analogous role by weighting poisoned versus benign supervision.
Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models · 2026Lomuscio, “Repairing misclassifica- tions in neural networks using limited data,” in Proceedings of the 37th ACM/SIGAPP Symposium on Applied Computing, 2022, pp.
Provable Fairness Repair for Deep Neural Networks · 2026The vulnerability of ML-based IoT systems to adversarial attacks - The lack of secure and robust ML models in the IoT context - The need for a comprehensive classification of AML and a systematic literature review of the latest research trends
Evaluation on three tabular benchmarks (Adult Census Income, German Credit, and Bank Marketing) reveals fairness concerns that vary widely across datasets: on Adult, 96.
FairLint-DL: An IDE-Native Tool for Fairness Debugging of Deep Learning Software · 2026In this paper, we first verify that existing mitigation mechanisms are insufficient for a new class of threats: indirect logic induction.
InceptionRAG: Stealthy Poisoning Attack Against Retrieval-Augmented Generation · 2026We close by mapping open challenges and arguing for standardized threat models and defenses that act proactively, in real time.
While previous studies have examined data poisoning, data drift and data integrity independently, limited research has integrated these governance risks within a single AI governance perspective.
Overall, our work unveils the pivotal yet understudied role of ICL in LLM safety, opening new avenues for understanding and improving them.
While substantial efforts have been dedicated to making neural networks invariant to 2D image translations and rotations, viewpoint invariance is rarely investigated.
Importance: Large language models (LLMs) are increasingly integrated into health care applications; however, their vulnerability to prompt-injection attacks (ie, maliciously crafted inputs that manipulate an LLM's behavior) capable of altering medical recommendations has not been systematically evaluated.
Vulnerability of Large Language Models to Prompt Injection When Providing Medical Advice · 2025 · DOIAlthough these initial results highlight NCCP as a theoretically sound and practically effective alternative to CP for uncertainty quantification with ACI in non-exchangeable scenarios, further empirical studies are warranted across diverse datasets and predictors.
In batch settings, inductive NCCP (INCCP) can outperform inductive CP (ICP) by utilising the full training dataset without requiring a separate calibration set, leading to improved efficiency, particularly when the data are limited.
An effective synergy of various defense mechanisms offers a promising approach to enhancing the robustness of ViT, yet this paradigm remains largely underexplored.
Hyper adversarial tuning for boosting adversarial robustness of pretrained large vision transformers · 2025 · DOIWe discuss the open challenges and research directions for foundation models in computer vision, including difficulties in their evaluations and benchmarking, gaps in their real-world understanding, limitations of contextual understanding, biases, vulnerability to adversarial attacks, and interpretability issues.
Despite their effectiveness, we still argue that these methods are under-explored in terms of determining how trustworthy the features are.
But beyond the weights, the overall structure and information flow in the network are explicitly determined by the neural architecture, which remains unexplored.
However, the potential of vision-language multimodal features for 3D mask presentation attack detection remains unexplored.
Second, potential of generated safety-critical scenarios to continuously improve ADS performance remains underexplored.
LLM-Attacker: Enhancing Closed-Loop Adversarial Scenario Generation for Autonomous Driving With Large Language Models · 2025 · DOIThe previous research has focused on designing backdoor attacks for CLMs, but effective defenses have not been adequately addressed.
Defending Code Language Models against Backdoor Attacks with Deceptive Cross-Entropy Loss · 2025 · DOIFinally, potential open issues and main research directions are highlighted for future consideration and enhancement.
While software-related concerns like adversarial attacks, private inference, and watermarking have been studied, the paper sheds light on previously underexplored hardware vulnerabilities such as trojans and side-channel attacks.
We finalize the discussion on challenges and open issues, as well as future research opportunities.
Backdoor Attacks to Deep Neural Networks: A Survey of the Literature, Challenges, and Future Research Directions · 2024 · DOILearning one policy that is both safe and robust under any adversaries remains a challenging open problem.
However, the vulnerability of DNN-based image ranking systems remains under-explored.
Most-cited papers in Adversarial Robustness in Machine Learning
- Membership Inference Attacks on Machine Learning: A Survey · ACM Computing Surveys · 2022 · 476 citations
- Explainable Artificial Intelligence Applications in Cyber Security: State-of-the-Art in Research · IEEE Access · 2022 · 399 citations
- Foundation Models Defining a New Era in Vision: A Survey and Outlook · IEEE Transactions on Pattern Analysis and Machine Intelligence · 2025 · 303 citations
- Explainable Artificial Intelligence in CyberSecurity: A Survey · IEEE Access · 2022 · 291 citations
- Making machine learning robust against adversarial inputs · Communications of the ACM · 2018 · 286 citations
- Dataset Security for Machine Learning: Data Poisoning, Backdoor Attacks, and Defenses · IEEE Transactions on Pattern Analysis and Machine Intelligence · 2022 · 267 citations
- Unlocking the black box: an in-depth review on interpretability, explainability, and reliability in deep learning · Neural Computing and Applications · 2024 · 207 citations
- The Smoke Detector Principle · Annals of the New York Academy of Sciences · 2001 · 187 citations
- Fast Yet Effective Machine Unlearning · IEEE Transactions on Neural Networks and Learning Systems · 2023 · 161 citations
- Adversarial Attack and Defense on Graph Data: A Survey · IEEE Transactions on Knowledge and Data Engineering · 2022 · 146 citations
Most recent work
- How malicious AI swarms can threaten democracy · Science · 2026
- Jailbreak and Guard Aligned Language Models With Only Few In-Context Demonstrations · IEEE Transactions on Pattern Analysis and Machine Intelligence · 2026
- From Edge Transformer to IoT Decisions: Offloaded Embeddings for Lightweight Intrusion Detection · Sensors · 2026
- Keeping Medical AI Healthy and Trustworthy: A Review of Detection and Correction Methods for System Degradation · IEEE Transactions on Biomedical Engineering · 2026
- Data Poisoning Vulnerabilities Across Health Care Artificial Intelligence Architectures: Analytical Security Framework and Defense Strategies · Journal of Medical Internet Research · 2026
- A survey of zero-knowledge proof based verifiable machine learning · Artificial Intelligence Review · 2026
- Security of Language Models for Code: A Systematic Literature Review · ACM Transactions on Software Engineering and Methodology · 2026
- Threats and vulnerabilities in artificial intelligence and agentic AI models · Frontiers in Artificial Intelligence · 2026
- Hiding in Plain Sight: RIS-Aided Target Obfuscation in ISAC · IEEE Transactions on Wireless Communications · 2026
- A survey of adversarial attacks on machine learning · Neurocomputing · 2026
Find a gap in your own Adversarial Robustness in Machine Learning sub-topic
This page shows what the Adversarial Robustness in Machine Learning literature already flags as unresolved. To narrow it to your specific question, search the Research Gap Finder: the search is free with a free account and lists the papers closest to your topic first. Unlocking that topic (50 credits, charged once) fills the comparison table from our 4.5M-paper local library and writes the gaps from its rows.
Open the Research Gap Finder →