Open research questions in Distributed systems and fault tolerance
65 unresolved questions extracted from the limitations and future-work sections of 295 Distributed systems and fault tolerance papers in our library. Each links back to the study that raised it.
What the literature leaves open
Prior research has not addressed the challenge of achieving high-performance transaction processing in geo-replicated OLTP databases - Prior research has not proposed a novel epoch-based asynchronous replication protocol
Ensuring transparency and independent validation of monitoring results. Eliminating the need for trust in a central authority. Achieving decentralized verification of Service Level Agreements.
The concurrent and physically distributed execution of callback functions. The lack of a full-fledged analysis of the multi-threaded behavior of ROS 2 callbacks and their scheduling. The subtle misconfigurations of ROS 2 applications.
The unknown players challenge, where the set of participants is unknown at the time of protocol deployment and is of unknown size. The player inactivity challenge, where participants can start or stop running the protocol at any time. The sybil challenge, where one participant may masquerade as many by using many identifiers.
The absence of bypass cannot be established from inside the governed path. There is a need for a method to measure the ungoverned surface instead of proving its absence.
Execution Coverage Reconciliation: measuring the ungoverned surface when absence of bypass cannot be proved from inside · 2026 · DOIInvestigating the application of FTI-TMR in various domains. Improving the stability scoring mechanism. Exploring the use of other consensus algorithms.
FTI-TMR: A Fault Tolerance and Isolation Algorithm for Interconnected Multicore Systems · 2026 · DOIExisting TMR variants have limitations in terms of energy overhead and fault coverage. There is a need for a fault-tolerant architecture that can address both transient and permanent faults.
FTI-TMR: A Fault Tolerance and Isolation Algorithm for Interconnected Multicore Systems · 2026 · DOIThe paper does not claim to solve epistemology, ontology, reference, meaning, or truth. Remaining open problems concern extensions, instantiation, refinement, normal forms, minimality, mechanization, and independent proof verification.
The Epistemic Map of Identity Persistence_ Regime Specification, Identity, Admissibility, Capacity, and Verification · 2026 · DOIThe relationships between the layers of the Identity-Persistence Program have not been fully understood. The program's architecture has not been fully mapped.
The Epistemic Map of Identity Persistence_ Regime Specification, Identity, Admissibility, Capacity, and Verification · 2026 · DOINo empirical run is attached; outcome status remains unperformed. The dated Lean PASS remains restricted to the preceding 2,777-claim surface; no new whole-corpus Lean PASS is claimed. The protocol does not assert a prototype, observed recurrence proxy, net dimensional power, efficiency, lossless switching or a successful experiment.
The paper identifies the need for a lawful apparatus test to measure the recurrence transition and both output channels. The preceding vacuum-beat protocol tests only the direct-repayment route.
Implementing the RSEP-XMSS instantiation and measuring its performance. Applying the VMC primitive to other stateful hash-based signature schemes. Investigating the use of VMC in other cryptographic protocols.
Verifiable Monotone Chains: A Primitive for Cryptographically Enforced State Lifecycles, with an Application to XMSS · 2026 · DOIThe lack of a way to make signer state discipline cryptographically verifiable in stateful hash-based signature schemes. The inability of verifiers to check for rollback attacks in stateful hash-based signature schemes. The need for a primitive that provides auditability and monotone-unforgeability security notions.
Verifiable Monotone Chains: A Primitive for Cryptographically Enforced State Lifecycles, with an Application to XMSS · 2026 · DOIHigh-stakes online assessment produces suspicious-event records, yet storage placement and burst admission remain insufficiently characterized.
A Kubernetes-Deployed Tamper-Evident Media-Evidence Provenance Pipeline with Hybrid Blockchain/IPFS Anchoring and Queue-Mediated Ingress · 2026 · DOIThe group dependence of NCI and C~ can be detrimental to their usefulness for detecting aberrant response patterns. There is a need for an index that is individually oriented and free from group dependence.
The paper plans to report on stability conditions for consensus functions at a later date.
The lack of a characterization of consensus functions for tree quasi-orders. The need for an analogue of Arrow's theorem for tree quasi-orders.
Future research is needed to determine the relationship o f independent measures o f client behavior (e.
The paper does not provide a comprehensive comparison with other Byzantine-resilient algorithms. The experimental evaluation is limited to a simulated environment. The paper does not discuss the computational complexity of FABA.
To explore the application of FABA to other machine learning algorithms. To evaluate the performance of FABA in real-world distributed environments. To design more robust Byzantine-resilient algorithms.
The paper suggests further research on the application of reflective parallel algorithms. The paper suggests further research on the development of programming languages and systems that support reflection.
The paper identifies a gap in the existing literature on behavioural theories of parallel algorithms. The paper addresses the need for a behavioural theory of reflective parallel algorithms.
The problem of storage access in autonomous AI systems, where storage errors are not just implementation concerns but part of the governance boundary. The need for a governed storage substrate that ensures storage safety and determinism in AI systems.
This architecture has several limitations. 1. No independent distributed‑consistency guarantee. VFS defines semantic storage boundaries. It does not, by itself, solve distributed transactions, object‑store consistency, or network partition be‑ havior. 2. Capability granularity and access‑control boundary. The initial VFS model intentionally exposes coarse‑grained capabilities such as read, write, delete, navigate, and query. Fine‑grained authorization, such as allowing writes only under a specific directory or only under a given user attribute, belongs to higher‑level policy evaluation. In this sense, ABAC and XACML‑style policy systems are treated as PDP/PEP‑layer mechanisms, while VFS provides truthful session capabilities and semantic storage facts consumed by those policies. 3. Provider honesty requirement. The model assumes providers honestly expose their capabili‑ ties. Provider identity contracts, signed manifests, testing, and attestation are needed for stronger assurance. 4. Semantic path design. The model depends on disciplined path and namespace design. Poorly de‑ signed namespaces can still produce confusing governance boundaries. 11. Relationship to Other Phase‑1 Papers This paper defines the physical and semantic storage boundary for AIKernel’s knowledge and data layer. 6 • Paper 01: ROM Format and Knowledge Snapshot Model defines the governed knowledge units stored and retrieved through VFS. • Paper 03: Pre‑Inference Admissibility Governance evaluates whether proposed storage operations are admissible before execution. • Paper 04: Trajectory Governance Model assumes storage operations are already bounded by VFS and governance policies. • Paper 05 / Paper 07: Execution and Implementation use VFS abstractions for execution state persis‑ tence, conversation snapshots, and replay archives.
Lack of admission control in LLM-based systems. Need for deterministic admission-control model.
Most-cited papers in Distributed systems and fault tolerance
- Secure multiparty computation · Communications of the ACM · 2020 · 239 citations
- Turbulences of speeding up data circulation. Frontex and its crooked temporalities of ‘real-time’ border control · Mobilities · 2020 · 20 citations
- Mechanistic Circuit-Breakers Generalize Across Irreversible Agent Actions and Architectures · Zenodo (CERN European Organization for Nuclear Research) · 2026 · 13 citations
- Graduated STA Control: Authority Governors and Runtime Deployment Architecture · Zenodo (CERN European Organization for Nuclear Research) · 2026 · 12 citations
- When Monitoring Facilitates Trust · Ethical Theory and Moral Practice · 2022 · 7 citations
- Bridging the gap between observation protocols and formative feedback · Journal of Mathematics Teacher Education · 2021 · 6 citations
- Consistency issue and related trade-offs in distributed replicated systems and databases: a review · RADIOELECTRONIC AND COMPUTER SYSTEMS · 2023 · 5 citations
- SAFE-Matter™ State-Machine Schema: Deterministic Runtime State Governance for Admissibility, Transition Integrity, and Execution Authority in Life-Critical Systems · Zenodo (CERN European Organization for Nuclear Research) · 2026 · 4 citations
- Log Replication in Raft vs Kafka · Studia Universitatis Babeș-Bolyai Informatica · 2020 · 4 citations
- Behavioural theory of reflective algorithms II: Reflective parallel algorithms · Science of Computer Programming · 2026 · 1 citations
Most recent work
- Mechanistic Circuit-Breakers Generalize Across Irreversible Agent Actions and Architectures · Zenodo (CERN European Organization for Nuclear Research) · 2026
- Graduated STA Control: Authority Governors and Runtime Deployment Architecture · Zenodo (CERN European Organization for Nuclear Research) · 2026
- SAFE-Matter™ State-Machine Schema: Deterministic Runtime State Governance for Admissibility, Transition Integrity, and Execution Authority in Life-Critical Systems · Zenodo (CERN European Organization for Nuclear Research) · 2026
- Behavioural theory of reflective algorithms II: Reflective parallel algorithms · Science of Computer Programming · 2026
- Improved Bounds for Coin Flipping, Leader Election, and Random Selection · 2026
- A High-Throughput Method for Fabric in Scenarios With Multiple Aborted Transactions · IEEE Transactions on Big Data · 2026
- Coalition Drift Mitigation: Why you cannot stop coalition drift — but you can contain it, observe it, and prevent it from becoming catastrophic. · Zenodo (CERN European Organization for Nuclear Research) · 2026
- K-means resilient to Byzantine faults · Progress in Artificial Intelligence · 2026
- Verifying wait-freedom for concurrent higher-order programs (Artifact) · Zenodo (CERN European Organization for Nuclear Research) · 2026
- Proof of Intent Consensus · Zenodo (CERN European Organization for Nuclear Research) · 2026
Find a gap in your own Distributed systems and fault tolerance sub-topic
This page shows what the Distributed systems and fault tolerance literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →