Computer Science · Research topic

Open research questions in Privacy-Preserving Technologies in Data

74 unresolved questions extracted from the limitations and future-work sections of 415 Privacy-Preserving Technologies in Data papers in our library. Each links back to the study that raised it.

What the literature leaves open

  • Task offloading in vehicular fog computing: State-of-the-art and open issues. Terahertz channel propagation phenomena, measurement techniques and modeling for 6G wireless communication applications: A survey, open challenges and future research directions. Federated optimization in heterogeneous networks: Advances and open problems. Outsourced and privacy-preserving collaborative k-prototype clustering for mixed data via additive secret sharing.

    Mobility-Aware and Privacy-Preserving Federated Reinforcement Learning with Multi-Paradigm Machine Learning for Edge Intelligence in 5G/6G Networks · 2026 · DOI
  • This paper introduced SCC-VFL, a server-centric framework for enforcing individual-level counterfactual stability in vertical federated learning while preserving predictive performance. The approach integrates three components: selective mask discovery that partitions each party’s features into non-descendants, mediators, and proxies; masked counterfactual generation that edits only policy-permitted mediators while preserving identity on non- descendants and suppressing proxy leakage; and a server-side selective counterfactual consistency loss that penalizes prediction changes under on-support mediator interventions. Across banking, healthcare, and criminal justice datasets, SCC-VFL achieves strong predictive utility while substantially reducing the selective consistency gap, flip rate, and adversarial attack success under both attribute inference and subspace-constrained PGD, in IID and non-IID vertical settings. Future work includes extending SCC-VFL along methodological, privacy, and deployment dimensions. An immediate direction is supporting higher-dimensional modalities such as text or imaging, and more complex vertical consortia with many parties, partial entity overlap, and asynchronous participation. On the privacy side, integrating formal differential privacy accounting and secure computation primitives into the mask discovery stage would provide end-to-end privacy guarantees beyond the current functional protections. Another avenue is adaptive and multi-attribute mask discovery that can track distribution shift and handle multiple protected attributes simultaneously. Finally, the masked generators developed for counterfactual stability could also support actionable recourse and auditing [45], where policy-compliant mediator edits are exposed to stakeholders as interpretable explanations or intervention recommendations. FAccT ’26, June 25–28, 2026, Montreal, QC, Canada Wasif et al.

    Toward Individual Fairness Without Centralized Data: Selective Counterfactual Consistency for Vertical Federated Learning · 2026 · DOI
  • Despite its advantages, implementation of FL in oncology faces challenges including data heterogeneity, model bias, communication limitations, cybersecurity risks, and lack of standardized validation frameworks.

    Federated Learning in Oncology: Privacy-Preserving Artificial Intelligence for Multi-Center Cancer Research · 2026 · DOI
  • Although the proposed framework shows promising results in combining Federated Learning (FL) with Explainable AI (XAI), there are still several areas where the system can be further improved. Future research can focus on enhancing performance, adaptability, and real-world applicability of the model. X.I. Personalized Federated Learning (pFL): One of the main challenges in federated learning is data heterogeneity, meaning that data collected from different hospitals may vary significantly in terms of quality, distribution, and patient demographics.In the current system, a single global model is shared across all clients. However, this approach may not always perform equally well for every institution. To address this issue, future work can explore Personalized Federated Learning (pFL).In this approach, instead of using one common model, each client can have a slightly customized version of the global model that better fits its local data. This can improve accuracy and make the system more adaptable to different healthcare environments. 554 International Journal of Advance and Innovative Research Volume 13, Issue 2: April - June 2026 ISSN 2394 - 7780 X.II. Integration of Multi-Modal Data: At present, the system mainly focuses on specific types of data such as medical images or structured health records. However, real-world healthcare data is much more complex and comes in different forms.

    FEDERATED AND EXPLAINABLE ARTIFICIAL INTELLIGENCE FOR PRIVACY-PRESERVING CLINICAL DECISION SUPPORT SYSTEMS · 2026 · DOI
  • A rigorous analysis of centroid and cluster-assignment stability under single-record perturbations remains an open problem. Finally, the effects of clipping and independent rounding on high-dimensional or strongly correlated data were not investigated and may lead to greater distortion than observed in the present study. In addition, the baseline comparison is limited to a non-private version and a coordinate-wise noise mechanism.

    Geometry-Based Differentially Private Synthetic Tabular Data Generation via K-Means Clustering with Bounded and Discrete Feature Constraints · 2026 · DOI
  • , from large neural networks); the current experiments used a modest three-layer network, so applicability to deep architectures remains to be verified. Fifth, the non-IID partitioning method (2–4 clusters per node) mimics realistic driving behavior differences, but real-world data may exhibit even higher heterogeneity, including completely missing classes for some nodes or extreme imbalance, requiring further investigation into the performance of weighted aggregation and anomaly detection under such conditions.

    Federated learning for privacy protection of connected autonomous vehicles in multi-access edge computing networks · 2026 · DOI
  • Niknam S, Dhillon HS, Reed JH (2020) Federated learning for wireless communications: motivation, opportunities, and 1 3Unfederated: Open Challenges, Deployment Gaps, and Emerging Directions in Federated Learning challenges. 33 90/s2 3177358 1 3Unfederated: Open Challenges, Deployment Gaps, and Emerging Directions in Federated Learning 82. Li X, Peng L, Wang Y-P, Zhang W (2025) Open challenges and opportunities in federated foundation models towards biomedical healthcare.

    Unfederated: Open Challenges, Deployment Gaps, and Emerging Directions in Federated Learning · 2026 · DOI
  • Existing PFL methods often face a fundamental tradeoff in which stronger global sharing can undermine local specialization, whereas stronger local adaptation can lead to overfitting under limited data, label imbalance, and missing class scenarios.

    Separate Aggregation of Split Network for Personalized Federated Learning · 2026
  • However, its practical effectiveness remains unclear, partly due to LLM pretraining, where overlaps and interdependencies with adaptation data can undermine privacy despite DP efforts.

    Benchmarking Empirical Privacy Protection for Adaptations of Large Language Models · 2026
  • This work presents FedMLAC, a unified federated audio classification framework designed to jointly address three central challenges in real-world federated learning (FL): data heterogeneity, model heterogeneity, and data poisoning. By introducing a lightweight, globally shared Plug-in model and a mutual learning mechanism between local and global models, FedMLAC decouples personalization from aggregation, facilitating effective cross-client knowledge transfer while supporting architectural flexibility. Moreover, to counteract the destabilizing effects of noisy or adversarial updates, a novel Layer-wise Pruning Aggregation (LPA) strategy is proposed, enhancing the resilience and stability of the global Plug-in model. Empirical evaluations across four diverse audio benchmarks confirm that FedMLAC consistently 26 outperforms existing FL baselines under clean and noisy conditions, both in homogeneous and heterogeneous model settings. Despite these advancements, several open challenges remain before prac- tical deployment. Ensuring efficient personalization in the presence of ex- treme client variability, minimizing communication overhead for low-resource edge devices, and maintaining robustness under adversarial manipulation or privacy-preserving constraints (e.g., differential privacy, secure aggregation) require further exploration. Addressing these general limitations will be crit- ical for building scalable and trustworthy FL systems for audio applications in increasingly complex, decentralized environments.

    FedMLAC: Mutual learning driven heterogeneous federated audio classification · 2026 · DOI
  • Conclusion The effective application of Privacy-Preserving Data Mining (PPDM) techniques is fundamentally challenged by the inherent trade-off between data privacy and data utility. While rigorous evaluation of this trade-off is essential for practical deployment, our review indicates that the current landscape of quality evaluation remains fragmented and predominantly reliant on static classification frameworks. To address these gaps, this study presented an extensive and structured consolidation of over 25 quality-level metrics across the entire data lifecycle, offering a holistic analysis that details their purposes, mathematical foundations, measurement focus, strengths, limitations, and specific applicability constraints. Beyond descriptive cataloging, we introduced a novel phase-based classification framework that aligns metric selection with the operational phases of PPDM: data collection, data publishing, and data mining. This taxonomy ensures that evaluation strategies are context-aware and tailored to the specific objectives and constraints of each processing stage. Furthermore, moving beyond isolated metric descriptions toward structured, operational guidance, we conducted a critical synthesis of inter-metric relationships—an analysis notably lacking in previous literature on PPDM. By systematically identifying orthogonal quality dimensions, complementary metric pairs, strategic substitution patterns, and hierarchical applicability structures, we demonstrated that single-metric evaluation is fundamentally inadequate for capturing the multidimensional nature of data quality in PPDM contexts. Building on these findings, this study proposed actionable, multi-metric evaluation protocols and decision frameworks—such as diagnostic matrices and tiered applicability hierarchies—which are designed to assist practitioners in developing robust, context-aware validation strategies. 123 192 Page 78 of 82 International Journal of Data Science and Analytics (2026) 22:192 Finally, we systematically examined the fundamental considerations and key challenges associated with measuring data quality within PPDM contexts. We explicitly identified the decisive factors that should inform appropriate metric selection, establishing a structured decision framework that offers practitioners strategic guidance for the selection and application of context-appropriate metrics. Ultimately, this work establishes a robust foundation for standardized, context-aware PPDM evaluation by providing comprehensive metric consolidation, introducing a practical phase-based classification framework, and conducting a systematic synthesis of inter-metric dynamics. The actionable protocols presented herein facilitate more rigorous, comparable, and reproducible assessments of privacy–utility trade-offs across various PPDM techniques, datasets, and application contexts.

    Quality metrics for Privacy-Preserving Data Mining: a systematic review, phase-based classification, and critical synthesis · 2026 · DOI
  • The DBI and Silhouette metrics provide complementary geometric perspectives on cluster quality through their distinct mathematical formulations and interpretational focuses. As internal metrics, they rely exclusively on the intrinsic properties of the sanitized data, emphasizing cluster compactness and separation. These metrics approach evaluation from fundamentally different viewpoints. As defined in equation (34), the DBI metric adopts a cluster-centric perspective, assessing the quality of the overall clustering structure. It evaluates each cluster by determining its worst-case similarity to a neighboring cluster and then averages these results. Conversely, as defined in equations (35-36), the Silhouette metric employs a point-centric perspective, evaluating the appropriate assignment of each individual data point. It compares a point’s average distance to its cluster members against its minimum average distance to points in other clusters. Consequently, the complementarity between these perspectives arises from their differing failure-detection capabilities. The DBI metric is proficient at identifying cluster-level pathologies, such as poorly separated or overlapping cluster boundaries. However, its cluster-averaging may obscure individual point misclassification (outliers) if the overall cluster geometry remains intact. In contrast, the Silhouette metric is effective at detecting point-level pathologies, such as outliers or misassigned individuals. Nevertheless, Silhouette’s point-averaging could mask systemic cluster-level issues if two adjacent, slightly overlapping clusters yield moderate scores because their points are closer to them than to more distant clusters. Based on this complementarity analysis, the proposed solution for a comprehensive assessment is a joint DBI- Silhouette analysis. Figure 17 depicts a 2x2 diagnostic matrix for the joint DBI-Silhouette analysis, utilized for assessing clustering results during the data mining phase. This combined approach enhances diagnostic capabilities, enabling evaluators to distinguish different scenarios, as follows: • Low DBI, High Silhouette: Ideal clustering—clusterlevel separation is excellent, and individual data points are appropriately assigned. • High DBI, Low Silhouette: Consistent poor quality—- cluster boundaries are unclear, and individual points are ambiguously assigned. • Low DBI, Moderate/Low Silhouette: Cluster structures are well-defined; however, individual point assignments may be suboptimal. • Moderate/High DBI, High Silhouette: Clusters exhibit some boundary ambiguity at the cluster level, yet individual points are confidently assigned to their nearest cluster. 123 It is important to note that both DBI and Silhouette share a common vulnerability to distance metric dependency. As both metrics rely on distance calculations, their values are sensitive to the choice of distance metric.

    Quality metrics for Privacy-Preserving Data Mining: a systematic review, phase-based classification, and critical synthesis · 2026 · DOI
  • A natural direction for future work is to extend our approach to cyclic patterns, which present additional challenges. Unlike acyclic patterns, cyclic patterns do not allow recursive decomposition, so their counts cannot always be derived by aggregating subpattern counts, making our multi-round aggregation strategy inapplicable. A possible workaround is to break cycles and reduce the problem to acyclic patterns with additional node constraints. For example, triangle counting can be reduced to counting 2-line paths whose endpoints are connected, but this introduces significant utility challenges, as it requires tracking both endpoints of each counted path while preserving privacy. Another promising extension is to support node-LDP, which offers stronger privacy guarantees than edge-LDP by protecting the presence of each individual node along with all its incident edges.

    Acyclic Graph Pattern Counting under Local Differential Privacy · 2026 · DOI
  • DP-SoftShape addresses the privacy-utility gap in time-series classification from a single angle: not every time step deserves the same amount of noise. A lightweight attention head scores patch importance and tilts the Laplace budget toward the patches that carry the class signal, while a downstream MoE block restores features that the noise still disrupts. Across 20 UCR datasets and four privacy budgets, this combination kept a higher mean accuracy than six standard LDP-input baselines, with the gap widening as the budget tightens. Two lines of follow-up work look most useful to us. The first is extending the patch-level budget split to multivariate streams, where channels interact and a naive split wastes budget on uninformative channels. The second is reducing the MoE block to a footprint suitable for on-device inference, since the current 2.7× overhead over ROCKET is the main cost we would like to remove. Tighter formal guarantees under the data-dependent allocation, and a hybrid with federated training, are also on our list.

    DP-SoftShape: Adaptive Differential Privacy via Attention-Guided Sparsification for Time-Series Classification · 2026 · DOI
  • DP-SoftShape has several limitations worth stating directly. Compute is the most visible one. The attention-driven saliency layer and the MoE refinement layer cost roughly 2.7× ROCKET and 7× InceptionTime in https://doi.org/10.53941/jmlis.2026.100009 11 of 14 Zhang et al. J. Mach. Learn. Inf. Secur. 2026, 2(2), 9 training time: averaged over the 20 × 5 = 100 (dataset, ϵ) configurations in our benchmark, mean per-configuration training time is 35.0 s (median 12.2 s, range 3.1 to 142.8 s), against 13.2 s for ROCKET, 4.8 s for InceptionTime, 3.6 s for Arsenal, and under 2 s for TSF, FCN and 1-NN-Euclidean. The cost is bounded but not free, and it matters most for the edge devices where local DP is most relevant. Beyond compute, the framework is hyperparameter-sensitive. How the global budget is split per-patch and how warm the expert routing temperature is set both affect accuracy across motif structures, and no single setting yet covers all 20 datasets without minor adjustment. The current scope is also univariate. Extending the patch-level budget split to multivariate streams under joint LDP, where channels interact and a uniform split wastes mass on uninformative channels, remains the most useful follow-up direction. On the privacy side, the guarantee here is per-patch ϵ/(1 − αn,p)-LDP under a data-dependent budget allocation, with the joint effective budget (cid:80) p ϵ/(1 − αn,p) depending on the input through the attention scores. A tighter worst-case bound on this data-dependent total budget, together with a data-independent ϵ-DP refinement of the mechanism, is left for future work.

    DP-SoftShape: Adaptive Differential Privacy via Attention-Guided Sparsification for Time-Series Classification · 2026 · DOI
  • successfully (FedAvg) algorithm system streaming allowing the system to generate a global model without compromising user privacy. Experimental results demonstrate that the system achieves high recommendation accuracy of around 94–96% while maintaining efficient response time. Furthermore, the system shows strong scalability and real-time performance, making it suitable for modern web applications such as e-commerce platforms, content services, and personalized web systems.

    A Privacy Preserving Personalized Search and Recommendation System Using Federated Learning and Web Usage Mining · 2026 · DOI
  • the accuracy. recommendation engine ranks the content items based on the combined scores derived from both local and global models. The system then provides top-ranked personalized recommendations to the user shown in figure 1.

    A Privacy Preserving Personalized Search and Recommendation System Using Federated Learning and Web Usage Mining · 2026 · DOI
  • preserving framework using federated learning and web usage mining techniques. Unlike traditional centralized systems, this approach ensures that user data remains on local devices, thereby enhancing data security and privacy. this system, user interaction data such as click frequency, browsing duration, scroll depth, and visit patterns are collected and stored locally on the client device. Instead of transferring raw data to a central server, a lightweight machine learning model is trained locally on each client using this interaction data. The model analyses user behaviour and computes preference scores for different content or websites.

    A Privacy Preserving Personalized Search and Recommendation System Using Federated Learning and Web Usage Mining · 2026 · DOI
  • This comprehensive survey has systematically examined the evolution of differential privacy algorithms over the past decade, providing a thorough analysis of their theoretical foundations, practical implementations, and performance characteristics. The research demonstrates that the PUXplore: Multidisciplinary Journal of Engineering , Vol 1 , Issue 1, Aug- Dec 2025|| Online (ISSN 3108-2106) 29 PUXplore: Multidisciplinary Journal of Engineering, Vol 1 , Issue 1, Aug – Dec 2025 || Online (ISSN 3108-2106) || DOI: https://doi.org/10.62373/yapwte32 field has matured significantly from its initial theoretical formulations to sophisticated algorithmic frameworks capable of addressing real-world privacy challenges across diverse application domains. The findings presented in this review offer valuable insights for researchers, practitioners, and policymakers involved in privacy-preserving data analysis. Synthesis of Research Objectives and Findings The study successfully addressed all six research objectives outlined in the introduction, yielding substantive findings that advance our understanding of differential privacy algorithms.

    Algorithmic Evolution of Differential Privacy: A Decade of Theoretical Advances and Practical Implementations · 2026 · DOI
  • processing, where large amounts of data are continuously processed. These models learn patterns from data and store them in their internal parameters. But, once trained, removing specific information from the model becomes difficult. This raises serious concerns related to privacy, security, and legal requirements such as the right to data deletion. Traditional methods, such as retraining significant the model after removing data, are effective but require time and computational resources. Other approaches, including fine-tuning and influence-based methods, try to reduce this cost but often fail to completely remove the effect of the targeted data.

    Machine Unlearning: Towards Privacy-Preserving and Trustworthy Artificial Intelligence System · 2026 · DOI
  • This paper presents a trusted heterogeneous FL scheme for edge scenarios, which has several notable strengths. First, it allows heterogeneous models from different end devices to contribute to the global model without requiring these devices to adopt uniform models, thereby maximizing data utilization across diverse devices. Second, the integrated distillation approach helps to reduce the burden on resource-constrained devices, allowing them to offload computationally intensive tasks to edge servers. Third, by incorporating a lightweight iterative masking mechanism, ELTFL ensures robust protection against privacy attacks, including gradient leakage, enhancing both privacy and security in FL. Despite its advantages, ELTFL has some limitations. One of the key challenges is that the effectiveness of KD largely depends on the quality and diversity of the small datasets used at the edge, which might not fully represent the data distribution of all participating devices. communication efficiency between the edge and cloud could still be improved, especially when scaling to large numbers of devices and datasets. Furthermore, In future research, we aim to investigate more efficient and adaptive KD techniques that can handle a wider range of heterogeneous models while further reducing communication costs. In addition, we plan to expand the scope of experiments to include larger and more diverse datasets, as well as real-world edge computing applications, to make ELTFL more robust in edge scenarios. Fig. 8 Attack success rate diagram for four schemes. the parameters of the global model are transmitted between edge servers and the cloud, the attacker is unable to restore the original image by gradient computation. In comparison, FedAvg exhibits a high ASR exceeding 99% between the edge cloud and the edge, indicating that FL without encryption or masking risks significant data leakage. In the case of attacking SMPC, the attack success rate is 0% due to the presence of double masks between the edges. However, since no mask is added between edge servers and cloud, the image can be restored. It is important to note that the image data obtained by the attacker are based on the aggregated gradient and may not necessarily reflect the original data. With PPFLEC, the edge data are not adequately protected and are directly aggregated at edge servers, making it susceptible to leakage. However, due to the inclusion of masks and hash functions between edge servers and the cloud, the data cannot be restored.

    ELTFL: A Trusted Heterogeneous Federated Learning in Edge Scenario · 2026 · DOI
  • Diagnosis support VisualDx Differential diagnosis via image matching. Standalone software; no real-time environmental awareness.

    Trustworthy intelligent rooms: integrating blockchain, federated learning, and data-centric AI for healthcare 4.0 · 2026 · DOI
  • SL addresses privacy and data integration issues, but research gaps exist, indicating potential areas for further exploration. • Security and Trust: While SL leverages blockchain for security and trust, further research is needed to address potential vulnerabilities, such as advanced cyber threats and insider attacks. Strong trust mechanisms and tailored security measures are crucial for SL networks. Swarm-FHE enhances SL security by integrating fully homomorphic encryption with blockchain, enabling secure collaborative model training even with compromised participants. Additionally, Li et al. combine blockchain and lightweight homomorphic encryption to ensure model security, data privacy, and computational efficiency, offering a competitive alternative to Federated Learning in remote ML applications. • Dynamic Node Management: Enhancing the robustness and dependability of SL systems may involve investigating dynamic techniques for node participation and incentive mechanisms to guarantee nodes’ continued and productive engagement in the swarm network. 24 Swarm Learning: A Survey of Concepts, Applications, and Trends A PREPRINT • Optimizing Leader Election: The leader election process in SL can lead to disproportionate bandwidth consumption, inefficiencies, and potential bottlenecks, causing dissatisfaction among participants and potentially compromising network security. To address these challenges, suggested refining the leader election mechanism for more equitable network load distribution. • Scalability and Efficiency: The ability of SL to expand across a growing number of nodes and a variety of data formats while maintaining efficiency and model performance should be investigated. Enhancing model aggregation techniques and communication protocols could be the main areas of research to facilitate widespread implementations of SL. • Interoperability and Standards: For SL to succeed, standards compliance and interoperability amongst various systems are essential. To solve issues with data format, protocols, and compliance, research could examine methods for SL to seamlessly integrate into existing IT systems. Qi et al. developed a blockchain twin mechanism to improve the interoperability and efficiency of SL on different blockchains, introducing an incentive mechanism for active participation, thus improving the overall performance and security of the SL process. • Energy Efficiency: Considering the possible magnitude of SL deployments, especially in the context of IoT, the development of power-saving learning algorithms is of the utmost importance. The emphasis of such research would be on minimizing the energy usage of devices involved in the SL process, a factor that is particularly critical for devices running on batteries or sensors located remotely.

    Swarm Learning: A Survey of Concepts, Applications, and Trends · 2026 · DOI
  • The ablation study demonstrates synergistic effects when combining federated learning, differential privacy, and homomorphic encryption, but does not investigate how the relative contributions of each component vary across different dataset sizes, feature dimensionality, or number of federated clients in the multi-layered privacy framework.

    Artificial Intelligence (AI) Based Multi-Layered Approaches for Privacy Preservation in Federated Learning · 2026 · DOI
  • The authors acknowledge applicability of their differential privacy and homomorphic encryption framework to finance and IoT domains but state that domain-level defenses are beyond their scope; specific privacy-utility tradeoffs and convergence characteristics for federated learning in these sensitive sectors remain unspecified.

    Artificial Intelligence (AI) Based Multi-Layered Approaches for Privacy Preservation in Federated Learning · 2026 · DOI

Most-cited papers in Privacy-Preserving Technologies in Data

Most recent work

Find a gap in your own Privacy-Preserving Technologies in Data sub-topic

This page shows what the Privacy-Preserving Technologies in Data literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.

Open the Research Gap Finder →

Related topics in Computer Science

74 open questions have been extracted from the limitations and future-work passages of 415 Privacy-Preserving Technologies in Data papers in our library. Each one below links back to the study that raised it, so you can read the original claim in context.

Tools for your next paper

Compare the categoryHonest roundups of the AI research tools, ours listed alongside the alternatives.

Command palette

Jump anywhere, run any action.