Computer Science · Research topic

Open research questions in Adversarial Robustness in Machine Learning

68 unresolved questions extracted from the limitations and future-work sections of 449 Adversarial Robustness in Machine Learning papers in our library. Each links back to the study that raised it.

What the literature leaves open

  • The vulnerability of Transformer-based NIDS to backdoor attacks and poisoning during training has not been characterized, despite recent work on backdoor detection in Transformers and the known susceptibility of NIDS to adversarial manipulation during both training and testing phases.

    Advancing Backdoor Attack Detection in Transformer Models using Feature Squeezing and Statistical Anomaly Filtering Techniques · 2026 · DOI
  • Knowledge Editing (KE) has emerged as a frontier for updating specific facts in LLMs without costly retraining, but its reliability and underlying mechanisms remain poorly understood.

    Exposing the Illusion of Erasure in Knowledge Editing for LLMs · 2026
  • AI systems are increasingly deployed for credit assessment and investment advisory in global financial markets, yet the integrity of their inference pipelines remains insufficiently addressed by existing regulatory frameworks.

    Invisible Manipulation Channels in AI-Assisted Financial Advisory: Implications for Market Integrity and Regulatory Design · 2026
  • Yet, pretrained LLMs already encode substantial harmful knowledge across many domains, which raises an important question: can an adversary jailbreak safeguarded models, to achieve harmful usage without fine-tuning at all? In this paper, we show that open-weight safeguards are susceptible to simpler strategies that, despite being well known, have not been systematically evaluated against these safeguards.

    Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks · 2026
  • Our findings highlight an important realism gap in current RAG security evaluation and suggest that poisoning in modern RAG systems should be studied as a multi-stage retrieval consistency problem rather than a retrieval-only problem.

    When Poison Fails After Retrieval: Revisiting Corpus Poisoning under Chunking and Reranking Pipelines · 2026
  • However, a notable limitation of the NFA is its tendency to misclassify benign neurons as mali- cious ones, which significantly degrades the primary task Page 19 of 21 296 accuracy of the global model.

    Neuron-level defense against backdoor attacks in federated learning · 2026 · DOI
  • This paper studied an emerging security risk: whether CodeLLMs can be structurally compressed while retaining the functionalitypreserving code mutation behavior that enables rapid generation of malware variants and evasion of static defenses. We introduced SecRL-Prune, an RL-based structured pruning framework that treats channel selection as a learned search problem using a KLguided teacher–student reward, together with a Top-𝑃 caching mechanism that reduces peak GPU memory by over 50% compared to prior methods. Across three real-world CodeLLMs and four compression ratios, SecRL-Prune consistently preserved more execution-level correctness (pass@k) and mutation diversity (var@𝑘) than a strong policy-learning baseline. These results show that code-mutation capability can survive significant structured compression, increasing the feasibility of compact mutation-focused models while helping defenders assess the risk posed by miniaturized mutation engines. Several limitations point toward future work: the reward is defined over a calibration distribution and truncated Top-𝑃 support, Figure 5: Peak GPU memory usage during pruning-policy training at 20% compression: SecRL-Prune consistently uses less than half the memory of PruneNet.

    SecRL-Prune: Structured Reinforcement Learning–Based Pruning of CodeLLMs for Preserving Adversarial Code Mutation · 2026 · DOI
  • In this paper, we introduced RESSAP, a robust ensemble framework that combines feature-level selection, data augmentation, and clas- sifier randomization to strengthen classifiers against adversarial evasion attacks. Our experimental evaluation shows that RESSAP improves robustness against adversarial evasion while maintaining strong accuracy on benign inputs. However, we acknowledge that the current evaluation is limited to a synthetically generated dataset, which may not fully capture the complexity of real-world applications. In addition, we have not yet compared RESSAP against alternative robust architectures that are specifically designed for adversarial settings. Future work will therefore focus on evaluating the proposed framework on diverse real-world datasets, conducting more exten- sive comparisons with state-of-the-art robust architectures, and refining the feature selection process to further improve adversarial robustness.

    Robust Ensemble of Selectively Strengthened and Augmented Predictors · 2026 · DOI
  • This review demonstrates that the failure of single-layer detectors is not model-specific but architectural in nature. Acoustic plausibility alone is insufficient to guarantee semantic validity, yet single-layer systems rely exclusively on this assumption. In contrast, multi- layered detection frameworks fundamentally alter the security landscape by integrating acoustic classification with linguistic verification through ASR systems. This layered approach enforces simultaneous acoustic and semantic constraints, transforming adversarial evasion into a significantly harder multi-objective problem. Both theoretical modeling and empirical evidence confirm that multi-layered architectures drastically reduce adversarial success across noise-based, reversed- audio, splicing, and GAN-based attacks. Linguistic verification restores critical phonetic and semantic information absent from feature-only detectors, while the combined decision boundaries reduce attack feasibility under diverse threat models and improve generalization across datasets and synthesis techniques. While challenges remain—including ASR robustness, multilingual handling, computational efficiency, and adaptation to future synthesis models—the evidence clearly establishes multi-layered frameworks as the most viable path forward. Effective audio deepfake defense must move beyond surface-level acoustic analysis toward integrated linguistic, semantic, and contextual validation. Such layered systems offer a foundation for secure, scalable, and trustworthy audio authentication capable of operating in increasingly adversarial environments.

    An Analytical Review of Multi-layered Security Frameworks for Audio Deepfake Detection: Combining ML Classifiers with Linguistic Verification · 2026 · DOI
  • Although multi-layered audio deepfake detection frameworks significantly improve robustness over in speech single-layer detectors, rapid advances synthesis and adversarial generation require continued evolution.

    An Analytical Review of Multi-layered Security Frameworks for Audio Deepfake Detection: Combining ML Classifiers with Linguistic Verification · 2026 · DOI
  • The evaluation covers English source text from one corpus of short snippets, and single-codepoint homoglyph substitution at swept but fixed rates. Multi-character visual confusions and confusions internal to the ASCII range, such as the letter l against the digit 1, are not measured. The per-source probe uses a single fixed carrier sentence; context-dependent normalization behaviour would not register. The evaluation does not include downstream task accuracy; XMR certifies that the model receives the same input as the clean case, not that its prediction is correct. uroman is evaluated on a 120-instance subsample for runtime, with the wide interval that n implies (CI 0.067 to 0.183 on R2). The attack is drawn from the same table that defines the tr39_skeleton baseline, which is why that row is reported as a ceiling rather than as a competing tool; a benchmark drawn from a different confusable inventory would move every number. translit is the author's tool, and the safeguard offered is reproducibility: the harness, corpus, seeds, and per-source results are all released.

    XMR: Exact Match Recovery for Evaluating Text Normalization Against Adversarial Unicode Attacks · 2026 · DOI
  • In this preliminary study, we empirically demonstrated that, through real-world experiments and Isaac Sim simulations, fingertip tac- tile sensors in embodied intelligent systems are vulnerable to EMI attacks. These results expose a previously underexplored attack surface in robotic perception and underscore a critical insight: en- suring the robustness of embodied intelligence requires a holistic view of sensor security, extending beyond vision to include tactile and other proprioceptive modalities. Building on these findings, several important scientific directions warrant further investigation. First, a systematic characterization of adversarial signal injection is needed to understand its fundamental capabilities and limitations, including attack range, power require- ments, temporal characteristics, alternative injection modalities, etc. Phantom Force: Injecting Adversarial Tactile Perceptions into Embodied Intelligence via EMI ASIA CCS ’26, June 1–5, 2026, Bangalore, India Extending this analysis across diverse tactile sensing technologies, such as capacitive, piezoresistive, optical, and multimodal sensors, will be essential to assess the generality of the threat and to identify modality-specific weaknesses. Second, future work should explore principled defense mechanisms that operate across the sensing and learning stack. Promising directions include hardware-level sensor design, signal-level anomaly detection and filtering, and learning-based approaches that improve robustness through sensor fusion, uncertainty modeling, and adversarially informed training. Ultimately, these efforts can contribute to a unified framework for trustworthy tactile perception, enabling embodied robots to operate in a secure and safe manner in adversarial noisy environments.

    POSTER: Phantom Force: Injecting Adversarial Tactile Perceptions into Embodied Intelligence via EMI · 2026 · DOI
  • This transformation reveals the core open problems—ground truth sources for ICS, initial acquisition of identity anchors, geometric structure of representation space—and provides clear directions for subsequent work. This points to a deep open problem: is our understanding of AI internal states sufficient to design reliable ICS and SV detectors? In this sense, OS-Net is not building a lock, but learning to stand—only a system that has grown bones can truly stand. Acquisition problem of identity anchors is unsolved: OS-Net assumes that the identity anchor IA already exists, but how IA is obtained—whether externally given or spontaneously generated by the system—is an unresolved open problem. 5 Open Problems The following open problems are left for future work: Precise computational complexity of NDC: Theorem 1 (Li, 2026b) proved the absolute invariance of NDC, but did not give an algorithm for computing NDC.

    Ontological Safety III: Why Incremental Approaches Fail, and How to Build an AI with Self-Boundaries · 2026 · DOI
  • Setting the measurement standard. This paper proposes a standard for how covert-channel egress is to be measured: pick a working encoder set, run each encoder against the defense, decode whatever survives, and report the Miller–Madow-corrected mutual information between embedded and recovered bits. We ship fifteen encoders as a realistic starting suite (Table 1): fourteen reach zero residual capacity and the mean-luminance encoder reaches a stated bound (Table 6). New encoders extend the suite directly — each stage is a separate extension, each encoder a benchmark row, and Section 5’s measurement applies unchanged. Which encoders to add is a concrete question: a production deployment runs in a known domain (finance, healthcare, legal, code review, content moderation), and the domain fixes the vocabulary, the payload shape, and the high-entropy fields an attacker can hide bits in. Attackers specialise to match: credential thieves work base64 and least-significant-bit (LSB) carriers, moderation evaders work homoglyphs and zero-width characters, audio exfiltrators work ultrasound. The threat surface that matters is the intersection of the deployment’s substrate and the encoders the relevant adversaries actually use; the benchmark should be extended to cover that. The floor of Section 2 remains: one bit always leaks through the agent’s choice to send or not send. 19 Text chokepoint sees text. The mediated hook exposes the outbound text of a message. Structured interactive blocks and channel-envelope metadata are not reachable at that point and need their own chokepoints. The media scramblers of Section 4 address the image and audio attachments, but they presuppose a media-egress chokepoint that the host must add. Bit estimates are upper bounds. The entropy scanner’s covert-bit estimates are deliberate over-counts. Over-counting covert capacity is the safe direction for a security control, but it means the ledger can be conservative and the scanner can charge benign high-entropy content (a legitimate base64 attachment, a random identifier). This is the false-positive surface the staged posture exists to manage, and it is why the entropy and behavioral stages audit before they enforce. Paced release has a latency cost. Constant-rate release delays messages. For an interactive agent this is a real user-visible cost, and it is why the strongest timing defense is configuration, not default. The LLM scrambler trusts a model. The model-backed rephraser sends outbound text to a language model, which adds three costs the deterministic stages do not carry: a latency and monetary cost, non-determinism, and a prompt-injection surface — the text being rephrased may itself contain instructions.

    An Application-Layer Multi-Modal Covert-Channel Reference Monitor for LLM Agent Egress. · 2026 · DOI
  • Four limitations scope the interpretation of the results: L1 - Synthetic data with high class separability. Experiments 1 and 2 used synthetic datasets where benign and adversarial classes are well-separated. The high accuracy metrics, therefore, demonstrate structural framework integrity rather than real-world generalization against sophisticated adversarial prompts. L2 - Proxy dataset for insider threat detection. Experiment 3 used the CICIDS-2017 network intrusion dataset as a proxy for litigation-pipeline insider threats. While CICIDS-2017 enables direct comparison with prior ZTA frameworks,, its traffic patterns do not capture the semantic characteristics of legal data exfiltration. L3 - Simulation-only validation. The framework was evaluated in a controlled simulation, not in a live legal AI production system. Production deployment would introduce additional variables, including concurrent user loads, heterogeneous LLM providers, and integration with existing document management infrastructure. L4 - Limited adversarial diversity. The 40-prompt Experiment 1 dataset, while carefully constructed, represents a narrow slice of the adversarial prompt landscape. Advanced multi-turn injection attacks, jailbreak chains, and indirect injections through retrieved documents in live RAG pipelines remain untested. FUTURE WORK Three directions will be implemented: (1) assessment of the prompt injection classifier on real attorney-written IP litigation queries; (2) implementation of ZT-IPLS in a live legal AI production system like Harvey or CoCounsel; and (3) formal regulatory preparedness certification to ABA and FRCP requirements using legal ontology systems.

    Zero-Trust Architecture Patterns for Securing AI-Driven IP Litigation Support Systems · 2026 · DOI
  • Three directions will be implemented: (1) assessment of the prompt injection classifier on real attorney-written IP litigation queries; (2) implementation of ZT-IPLS in a live legal AI production system like Harvey or CoCounsel; and (3) formal regulatory preparedness certification to ABA and FRCP requirements using legal ontology systems.

    Zero-Trust Architecture Patterns for Securing AI-Driven IP Litigation Support Systems · 2026 · DOI
  • Description Robot-assisted surgery and intervention Preventive, Corrective, Normative Pre-action approval; shared control; emergency stop High-risk maneuvers; proximity to sensitive anatomy Reduced procedural risk; increased clinician trust Reduced autonomy; workflow disruption; cognitive load Source…

    Humans as Safety Constraints: A Survey of Human-in-the-Loop Reinforcement Learning for Critical Systems · 2026 · DOI
  • Reaction latency; trust miscalibration; operator [66–67] fatigue Source: Author ©2026, Cognizance Journal, ZAIN Publications, Fridhemsgatan 62, 112 46, Stockholm, Sweden, All Rights Reserved 27 Kenneth Besigomwe, Cognizance Journal of Multidisciplinary Studies, Vol.6, Issue.4, April 2026, pg. 15-37 (An Open Accessible, Multidisciplinary, Fully Refereed and Peer Reviewed Journal) cognizancejournal.com ISSN: 0976-7797 Impact Factor: 5.503 Index Copernicus Value (ICV) = 92.57 Figure 6’s radar chart visualizes key human limitations, showing that latency and cognitive load are the dominant factors constraining intervention effectiveness. While safety drivers prevent many critical incidents, reliance on sustained human vigilance introduces cognitive strain and reaction time delays. These limitations motivate event-triggered or confidence-aware engagement strategies, where human input is prioritized for high-risk events, consistent with the Human Safety Constraint Framework (HSCF). The table and narrative together illustrate how human authority is embedded structurally rather than incidentally in deployed systems. 4.7.3 Medical Robotics Medical robotics involves irreversible consequences in the event of failure, stringent ethical standards, and close human–robot collaboration. Reviews of robot-assisted surgery consistently report low rates of major complications, typically on the order of a few percent or less, particularly in mature procedures and high-volume centers. Nevertheless, the potential severity of rare failures necessitates conservative human oversight and deliberate authority allocation in safety-critical surgical workflows [(68-70)]. Human safety constraints in medical robotics operate across preventive, corrective, and normative roles. Preventive constraints include pre-action authorization, safety envelopes, and shared control schemes that restrict robot motion near critical anatomy. Corrective constraints rely on emergency stop mechanisms and immediate clinician override. Normative constraints enforce adherence to medical ethics, professional accountability, and standards of care. Table 2 summarizes the human safety constraints, enforcement mechanisms, and observed outcomes for medical robotics systems, highlighting preventive, corrective, and normative roles based on the surveyed studies [68–71].

    Humans as Safety Constraints: A Survey of Human-in-the-Loop Reinforcement Learning for Critical Systems · 2026 · DOI
  • CatBoost achieved macro-F1 of 0.98918 and 0.99156 accuracy on SmartBugs-Wild cluster classification, substantially outperforming vulnerability detection (F1 ~0.78-0.79). The transferability of cluster-level structural representations to vulnerability detection tasks within the same blockchain-integrated framework has not been explored.

    Blockchain-integrated machine learning framework for transparent smart contract vulnerability detection · 2026 · DOI
  • The confusion matrices show that XGBoost commits errors between neighboring vulnerability classes (denial_of_service and arithmetic) more frequently than CatBoost and Random Forest, but the feature space geometry and decision boundary characteristics causing these misclassifications in smart contract vulnerability detection have not been examined. Analyzing the geometric properties of feature representations for confused vulnerability pairs would clarify this phenomenon.

    Blockchain-integrated machine learning framework for transparent smart contract vulnerability detection · 2026 · DOI
  • Random Forest demonstrated inherent robustness with nearly identical performance before and after SHAP application, but the specific ensemble-based averaging mechanisms that provide this interpretability-independent robustness on smart contract vulnerability detection have not been formally analyzed. This requires theoretical investigation of why tree ensemble averaging outperforms gradient boosting for unbalanced vulnerability classification.

    Blockchain-integrated machine learning framework for transparent smart contract vulnerability detection · 2026 · DOI
  • The SmartBugs-Wild dataset evaluation revealed highly skewed cluster distributions (35,499 contracts in Cluster 1 vs. 149 in Cluster 2) with only 34.15% variance explained by the first two principal components, yet the impact of this structural heterogeneity on model generalization across clusters has not been evaluated. Cross-cluster validation performance between models trained on different structural archetypes of smart contracts should be investigated.

    Blockchain-integrated machine learning framework for transparent smart contract vulnerability detection · 2026 · DOI
  • XGBoost and LightGBM showed lower recall (~0.82) compared to Random Forest and CatBoost (~0.77-0.88) on smart contract vulnerability detection due to overfitting on dominant features and sensitivity to hyperparameter tuning. The specific hyperparameter configurations that minimize this performance gap for gradient boosting methods on blockchain vulnerability datasets have not been systematically explored.

    Blockchain-integrated machine learning framework for transparent smart contract vulnerability detection · 2026 · DOI
  • The SHAP-based optimization framework improved CatBoost's minority class detection (arithmetic and denial_of_service vulnerabilities) but the mechanism by which SHAP restructures feature importances to reduce off-diagonal errors in imbalanced smart contract datasets remains theoretically unexplained. Future work should investigate whether this improvement generalizes to other vulnerability types or datasets with different imbalance ratios.

    Blockchain-integrated machine learning framework for transparent smart contract vulnerability detection · 2026 · DOI
  • Limitations of this study include the discrete nature of the action space (on/off gating) and the potential for over- head from the DRL agent itself, although it was observed Third, closer examination is warranted of the dynamics of the control loop itself.

    Hard Real-Time Deep Learning for Security-Critical Streams: An Integrative Algorithmic Framework for Cybersecurity, Scientific Analytics, and Decision Making · 2026 · DOI

Most-cited papers in Adversarial Robustness in Machine Learning

Most recent work

Find a gap in your own Adversarial Robustness in Machine Learning sub-topic

This page shows what the Adversarial Robustness in Machine Learning literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.

Open the Research Gap Finder →

Related topics in Computer Science

68 open questions have been extracted from the limitations and future-work passages of 449 Adversarial Robustness in Machine Learning papers in our library. Each one below links back to the study that raised it, so you can read the original claim in context.

Tools for your next paper

Compare the categoryHonest roundups of the AI research tools, ours listed alongside the alternatives.

Command palette

Jump anywhere, run any action.