Open research questions in Scientific Computing and Data Management
184 unresolved questions extracted from the limitations and future-work sections of 557 Scientific Computing and Data Management papers in our library. Each links back to the study that raised it.
What the literature leaves open
Existing solutions decouple data and model provenance. Existing solutions provide non-tamper-evident security. Existing solutions use trust assumptions for central entities. Existing solutions are unscalable for IoT-sized data.
Tamper-Evident Data and Model Provenance for IoT-Based Machine Learning Using Blockchain and Off-Chain Storage · 2026 · DOICan we trust agentic reviewers for agent-generated papers? SAR and PR both over-credit agent- generated papers relative to human reviewers (§5.1). Future automated reviewers should combine with principled calibration against human review. Faithfulness over complex tasks. Frontier model providers increasingly advertise faithfulness as a core capability of their agents, yet under our open-ended end-to-end research setting we still observe substantial fabrication (§5.2). The claim of faithful behaviour does not yet survive contact with sufficiently complex tasks. Future work should focus on training agents to be faithful end-to-end rather than only on individual reasoning traces. Better experiment-planning agents and scaffolds. The major challenge for auto-research is experimental rigor (§5.2). Closing this gap will require improvements: stronger agent capabilities and scaffolds that harness agents for designing and executing rigorous experiments end-to-end.
How Far Are We From True Auto-Research? · 2026Further research is needed to address the limitations of ProHunter, - Future work could focus on improving the accuracy and efficiency of ProHunter, - Research on integrating ProHunter with other threat hunting solutions could be beneficial
Significant memory and time overhead due to the extremely large provenance graphs. Imprecise segmentation of APT activities from provenance graphs due to their intricate entanglement with benign operations. Poor alignment of attack representations between CTI-derived query graphs and provenance graphs due to their substantial semantic gaps.
The lack of guidance on designing efficient, scalable, and reproducible AI-driven HPC workflows. The need for a framework to transition from rigid execution pipelines to adaptive, intelligent computational environments. The challenge of managing heterogeneous resources and complex workflow orchestration in AI-driven HPC workflows.
Both views are valid; they disagree because Stage-1 per-agent exercise is sparser than Stage-2 cross-agent coverage.
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows · 2026Synthetic perturbations appear to offer inexpensive calibration data for LLM evaluators in biomedical ML, where expert review is scarce.
Four Ledgers, Not One Score: Responsible Communication of LLM-Judge Calibration in Biomedical ML · 2026Environment specification and full rerunnability remain weak for every evaluated system, ours included, so reproducibility is still an open problem for scientific agents.
OpenAI4S: Code as Action, Science as Sessions · 2026These assumptions are often implicit and rarely examined, even though results can depend strongly on the synthetic model being used.
Cophylogeny simulators are not interchangeable: similarities, differences and structural biases in synthetic host-symbiont · 2026 · DOIThe determinism claim is untested: the specification claims two architects given one rule set produce one assignment, and that has not been tested with two architects — a procedure can be deterministic on paper and still admit judgement where a workload is assigned to a class.
Cross-Border AI Architecture Patterns: Three Patterns, and Why the Fourth Is Not a Choice · 2026 · DOIThe determinism claim is untested: the specification claims two architects given one rule set produce one assignment, and that has not been tested with two architects — a procedure can be deterministic on paper and still admit judgement where a workload is assigned to a class.
Cross-Border AI Architecture Patterns: Three Patterns, and Why the Fourth Is Not a Choice · 2026 · DOISection 13 is explicit about what a reference architecture does not establish — there is no implementation and no evaluation here — and about the open problems that remain.
The Knowledge Layer: A Reference Architecture for Delegated AI Action in Regulated Institutions · 2026 · DOIDespite their potential influence, the specific contributions of STEM professionals in AI policy remain underexplored.
The paper concludes that future research should focus on improving model efficiency, enabling multimodal data fusion, and enhancing compatibility with existing industry tools to realize the full potential of LLMs and support the digital transformation of the AECO sector.
Application of Large Language Models in the AECO Industry: Core Technologies, Application Scenarios, and Research Challenges · 2025 · DOIHowever, our laboratory routines and our scientific practice of communicating scientific results are insufficient to ensure the reproducibility and scalability of experiments, and data management has become a bottleneck to progress in biocatalysis.
The need for a comprehensive testing and validation framework for EOSC Beyond Pilot Nodes. The gap in establishing a federated European Open Science Cloud infrastructure.
HiRSE Seminar | Metadata for Research Software | M.Gruenpeter | 25/03/2026 | CC-BY 4.0 | #25/46 Objective: #RSMD_cheklist To ensure the collection, curation, and maintenance of research software metadata, the following general requirements are recommended for end users, including researchers, software engineers, curators, and institution staff.
Making Research Software Visible, Citable, and Preserved: A Metadata Deep Dive for RSEs · 2026 · DOIInappropriate or superficial use of containers can undermine their benefits. There is a need for practical tips on using software containers effectively in computational biology research.
The complexity of writing parallel applications is a significant challenge. The need for a user-friendly interface for managing distributed, task-based workflows is identified.
DistributedWorkflows.jl - A Julia interface to a task-based workflow management system. · 2026 · DOIThis paper and the DistributedWorkflows.jl package intro- duce a user-friendly interface for managing distributed, task-based workflows. It enables users to generate, visualize, compile, and launch workflows represented as Petri nets through simple meth- ods. The package provides a fully documented public API, allow- ing users to locally test applications before deploying them on ex- pensive clusters, ensuring cost efficiency. With binaries available for multiple Linux distributions, the tool is currently best suited for long-running processes. Upcoming updates aim to enhance the package with specialized transition types for reduced boilerplate code, additional workflow examples to guide users, and convenience functions that streamline Fig. 6. Part of a Petri net modeling different repair synthesis pathways of DNA as in [26]. the process of creating and managing workflows. Furthermore, im- provements to the user interface are planned, making the tool even more intuitive. These features are expected to expand the package’s utility, enabling broader adoption and more efficient handling of distributed workflows across various applications. 8. ACKNOWLEDGEMENTS We extend our gratitude to Fraunhofer ITWM for funding the Ph.D. studies of the first author. As well as, enabling her to work on and develop DistributedWorkflows.jl. We are also deeply grate- ful to Prof. Dr. Anne Frühbis-Krüger and Prof. Dr. Claus Fieker for their invaluable insights into the needs of domain scientists, which guided the design of this package with a strong focus on user experience. Our sincere thanks goes to Dr. Tiberiu Rotaru for his expertise in GPI-Space, particularly during the initial stages of integrating our Julia prototype with C++.
DistributedWorkflows.jl - A Julia interface to a task-based workflow management system. · 2026 · DOIThe widespread over-allocation of resources leads to increased queue wait times, costs, and carbon emissions. There is a need for automatic right-sizing and efficient scheduling of job resource requests within a science gateway.
Right-Sizing Compute Resource Allocations for Bioinformatics Tools with Total Perspective Vortex · 2026 · DOILooking ahead, two directions stand out. First, the automation of the shared database lifecycle remains a key priority. While defaults curated by experienced administrators and validated against production data have proven highly reliable, scaling to thousands of tools will require semi-automated updates. Incremental approaches, such as bracketing input sizes for tools with predictable scaling or selectively benchmarking high-impact applications, could provide efficient pathways toward automation without overreliance on machine learning. Second, there is potential for closer integration with live system telemetry. The meta-scheduling function at Galaxy Australia already queries backend load in real time, but similar approaches could extend to memory or I/O characteristics. Combining static defaults with live data would yield more adaptive scheduling while maintaining interpretability. Third, the shared TPV configuration database, combined with historical job metrics, could serve as training data for ML models that infer baseline configurations for new tools lacking curated entries. This is particularly appealing for the cold-start scenario where a new tool is added and no prior execution data exists. Such a hybrid approach could complement the rule-based system by providing informed starting points while retaining TPV’s interpretability and administrator override capabilities.
Right-Sizing Compute Resource Allocations for Bioinformatics Tools with Total Perspective Vortex · 2026 · DOIThe paper does not provide a comprehensive comparison with all existing early exiting methods. The evaluation is limited to twelve datasets and eight foundation models.
Evaluating EPEE on more datasets and foundation models. Investigating the application of EPEE in other domains beyond biomedicine. Developing more efficient and effective early exiting methods.
Despite this expansion, there exists no standardized, cryptographically verifiable infrastructure capable of proving the provenance, reproducibility, model origin, training lineage, or verification status of AI-generated artifacts.
INTERGET v2.0: The Provenance and Autonomous Coordination Layer of the Sovereign Web — Intelligent Network for Trusted Execution of Recorded Generative Experiences and Trajectories · 2026 · DOI
Most-cited papers in Scientific Computing and Data Management
- Science mapping software tools: Review, analysis, and cooperative study among tools · Journal of the American Society for Information Science and Technology · 2011 · 2,759 citations
- Fab · Communications of the ACM · 1997 · 2,037 citations
- SciMAT: A new science mapping analysis software tool · Journal of the American Society for Information Science and Technology · 2012 · 1,068 citations
- The Galaxy platform for accessible, reproducible, and collaborative data analyses: 2024 update · Nucleic Acids Research · 2024 · 851 citations
- Examining the Challenges of Scientific Workflows · Computer · 2007 · 386 citations
- Implementing faceted classification for software reuse · Communications of the ACM · 1991 · 319 citations
- Proactive computing · Communications of the ACM · 2000 · 289 citations
- From the Semantic Web to social machines: A research challenge for AI on the World Wide Web · Artificial Intelligence · 2009 · 181 citations
- The 2012 free and open source GIS software map – A guide to facilitate research, development, and adoption · Computers Environment and Urban Systems · 2012 · 176 citations
- BindingDB in 2024: a FAIR knowledgebase of protein-small molecule binding data · Nucleic Acids Research · 2024 · 168 citations
Most recent work
- Provenance Erasure Rate: A Compression-Survival Metric for Attribution Loss in AI-Composed Search Outputs · Zenodo (CERN European Organization for Nuclear Research) · 2026
- Erasure Skew: A Measurement Program for the Power-Conditioning of Provenance Loss in Retrieval and Composition Systems · Zenodo (CERN European Organization for Nuclear Research) · 2026
- Control and Reliability for Research Continuation: Observability, Failure Propagation, and Validation in Transcript-Sufficient Systems · Open MIND · 2026
- Constitutive Provenance: A Sentence-Level Countermeasure Against Composition-Layer Erasure · Zenodo (CERN European Organization for Nuclear Research) · 2026
- Provenance After AI: Semantic Provenance and the Provenance Erasure Rate as Extension of C2PA, Data Provenance Initiative, and EU AI Act Frameworks (v1.1) · Zenodo (CERN European Organization for Nuclear Research) · 2026
- Formal Foundations of Semantic Physics (EA-SEI-FF-01, v0.2 — Post-Assembly Review) · Zenodo (CERN European Organization for Nuclear Research) · 2026
- Provenance Is What Authorship Must Endure: AI-Mediated Writing, the Authorship-Slop Distinction, and the Missing Third Dimension of Provenance · Zenodo (CERN European Organization for Nuclear Research) · 2026
- Integrity Lock Certificate: The Refutation Triad — Mutual Anchoring of the AI_Bleeding Refutation Dossier (EA-LOCK-AIBLEEDING-01 v1.0) · Zenodo (CERN European Organization for Nuclear Research) · 2026
- KBase: Open-source Platform for Collaborative Biological Data Analysis and Publication · Journal of Molecular Biology · 2026
- ADR-001: Genesis ODE — Core vs Experimental Scope · Zenodo (CERN European Organization for Nuclear Research) · 2026
Find a gap in your own Scientific Computing and Data Management sub-topic
This page shows what the Scientific Computing and Data Management literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →