Open research questions in Data Quality and Management
96 unresolved questions extracted from the limitations and future-work sections of 501 Data Quality and Management papers in our library. Each links back to the study that raised it.
What the literature leaves open
We built a NL2SQL system for enterprise use that uses a multi-agent orchestration over enriched schema metadata with retrieval, planning, reflection, and guarded execution. The approach handles large, complex schemas by finding the most relevant tables and columns for a given query. Across the 11-domain subsets of the BIRD benchmark dataset, the system consistently outperforms non-agentic approach. Vector search and schema enrichment raise semantic accuracy more than exact match and reduce number of query iterations resulting in low latency. The agent is able to attain over 78.1% semantic accuracy on BIRD dataset. Further improvement avenues include extending the agent to handle various SQL dialects and domains, and enhancing search and ranking for more accurate retrieval.
AgentNLQ: A General-Purpose Agent for Natural Language to SQL · 2026The integration of semantic interpretation, planning, tool use, execution, validation, and result synthesis within a unified process introduces new risks and challenges. The need to distinguish different levels of autonomy and characterize autonomy as a continuum determined by the degree of user control, the actions an agent is authorized to perform, and the oversight or approval mechanisms required. The challenge of providing a comprehensive framework that integrates and extends existing foundations to provide essential capabilities for trustworthy agent-mediated data processes.
From Data Quality to Quality of Agentic Data Use: A Conceptual Framework for Agentic Data Engineering · 2026 · DOIThe existing foundations such as Data Contracts, Semantic Layers, Data Quality, Guardrails, AI Governance, and Data Provenance provide fragmented capabilities. There is a need for a comprehensive framework that integrates and extends these foundations to provide essential capabilities for trustworthy agent-mediated data processes. The current approaches do not fully address the risks introduced by the shift of data-engineering automation toward end-to-end processes.
From Data Quality to Quality of Agentic Data Use: A Conceptual Framework for Agentic Data Engineering · 2026 · DOIExisting data cleaning tools use static, rule-based approaches and do not accommodate the unique needs of a specific domain. Data cleaning is one of the most time-consuming phases of the entire data lifecycle.
Cleansera: A Context-Aware, Algorithm-Centric Data Cleaning System with RAG-Enhanced Intelligence · 2026 · DOIThe paper states future work will evaluate Cleansera's performance and empirical efficacy across industry datasets, but does not specify which industry types (healthcare, finance, e-commerce, manufacturing), dataset sizes, data quality profiles, or cleaning rule complexities will be tested. Comparative performance evaluation against existing data cleaning systems on standardized benchmarks is not outlined.
Cleansera: A Context-Aware, Algorithm-Centric Data Cleaning System with RAG-Enhanced Intelligence · 2026 · DOIThe gap between AI prototypes and enterprise AI motivates the study. Regulated industries face a structural constraint in deploying AI due to the need for auditability, traceability, and lifecycle risk controls.
Data Architecture Maturity as A Predictor of Enterprise AI Success in Regulated Industries · 2026 · DOIThe integration of heterogeneous data sources remains a complex challenge due to schema diversity, semantic inconsistencies, and data quality issues. Conventional rule-based data integration pipelines often struggle to handle schema heterogeneity, semantic inconsistencies, and incomplete records.
An Artificial Intelligence-Based Data Integration Framework for Real-Time Cross-Source Data Harmonization · 2026 · DOIManaging dual temporal dimensions in relational systems introduces performance overhead. Retroactive updates require cascading interval adjustments across historical records. There is a need for hybrid DBMS infrastructures.
Comprehensive insights into bitemporal databases: a PRISMA-guided systematic literature review · 2026 · DOIThe need for hybrid DBMS infrastructures. Limited empirical validation on industrial-scale datasets. Insufficient cloud-native temporal optimisation strategies.
Comprehensive insights into bitemporal databases: a PRISMA-guided systematic literature review · 2026 · DOIData heterogeneity. Spatial complexity. The need to comply with the FAIR principles.
A database-driven research data framework for integrating and processing high-dimensional geoscientific data · 2026 · DOIFurther development of the framework to support large-scale research projects. Integration of the framework with other data management systems. Application of the framework to other domains, such as environmental science or biology.
A database-driven research data framework for integrating and processing high-dimensional geoscientific data · 2026 · DOIThe cost of pairwise execution of ML models is prohibitive. Existing algorithms have limitations, including lack of statistical guarantees or inefficiency.
Accelerating Approximate Analytical Join Queries over Unstructured Data with Statistical Guarantees · 2026 · DOIExisting algorithms either fail to provide statistical guarantees or become as inefficient as uniform sampling. There is a need for a method that simultaneously achieves statistical guarantees and high efficiency.
Accelerating Approximate Analytical Join Queries over Unstructured Data with Statistical Guarantees · 2026 · DOIThe paper identifies a need for guidance on standardizing passport data for publication in Genesys. The MCPD standard requires specific formatting for data such as dates and geographical coordinates.
The need for a practical guide on passport data preparation using Excel. The lack of standardization and validation in genebanks managing passport data in spreadsheets.
The infodemic has far-reaching societal, political, and economic impacts. Current detection, prevention, and mitigation strategies remain fragmented and reactive. The problem of information quality assessment is understudied.
The role of diversity was less stable in the academic setting, offering only marginal gains and sometimes diminishing at higher levels of dispersion. The model was demonstrated in the context of academic publications, and its applicability to other domains is unknown.
The many single-application studies proposing hundreds of dimensions make data quality assurance a daunting prospect. The diversification of industries is compounding this challenge, making a multi-sector classification essential.
The DaTUM framework: a multi-sector thematic analysis of data quality dimensions and their impacting factors · 2026 · DOIExisting frameworks treat data quality dimensions as context-independent and universally applicable, yet no systematic method exists for surfacing and interrogating the latent value commitments and normative assumptions embedded in quality assessment frameworks themselves. Prior work identifies multiple dimensions (accuracy, completeness, etc.) but does not examine how the selection and weighting of these dimensions encodes particular conceptual frameworks and priorities that may systematically distort phenomena when applied across different use contexts.
The DaTUM framework: a multi-sector thematic analysis of data quality dimensions and their impacting factors · 2026 · DOIThe lack of integration of employee databases, daily activity reporting, and operational monitoring within a unified information system. The reliance on manual or spreadsheet-based reporting mechanisms that are vulnerable to data redundancy and inconsistent updates.
Design and Development of an Employee Database and Daily Reporting Information System Using Rapid Application Development (RAD) · 2026 · DOIThe complexity of the battery cell production process chain. The need for cost-efficient quality control measures. The requirement for real-time monitoring of key performance indicators.
Enabling Holistic Tracking and Tracing in Battery Cell Production: Data Management and Applications · 2026 · DOIThe lack of a comprehensive tracking and tracing system in battery cell production. The need for data-driven approaches to navigate the complexities of battery cell production. The requirement for cost-efficient quality control measures in battery cell production.
Enabling Holistic Tracking and Tracing in Battery Cell Production: Data Management and Applications · 2026 · DOIThe lack of effective methods to prevent and correct integrity constraint violations in relational databases. The need for a robust solution to maintain data quality and consistency in digitalized systems.
Integrating AI and Advanced Algorithms for Sustainable Data Integrity in Digitalised Systems · 2026 · DOIDeveloping dynamic risk maps for zoonotic disease spillover using integrated livestock, biodiversity, and public health data. Investigating the application of AI-driven harmonisation to other domains where data integration is challenging. Exploring the potential of the federated data network to inform evidence-based decision-making for managing the planet's interconnected agricultural and natural ecosystems.
Connecting the herd and the habitat: A plea for an AI-driven framework for integrating livestock and biodiversity data · 2026 · DOIThe current data divide between agricultural, public health, and biodiversity sectors is a primary barrier to implementing a true One Health approach. The lack of integrated data on livestock and biodiversity hinders our understanding of the ecological and public health impacts of global livestock production.
Connecting the herd and the habitat: A plea for an AI-driven framework for integrating livestock and biodiversity data · 2026 · DOI
Most-cited papers in Data Quality and Management
- A relational model of data for large shared data banks · Communications of the ACM · 1970 · 4,978 citations
- eTailQ: dimensionalizing, measuring and predicting etail quality · Journal of Retailing · 2003 · 1,528 citations
- Mediators in the architecture of future information systems · Computer · 1992 · 1,282 citations
- The KDD process for extracting useful knowledge from volumes of data · Communications of the ACM · 1996 · 1,141 citations
- Data governance: Organizing data for trustworthy Artificial Intelligence · Government Information Quarterly · 2020 · 582 citations
- The Role of ChatGPT in Data Science: How AI-Assisted Conversational Interfaces Are Revolutionizing the Field · Big Data and Cognitive Computing · 2023 · 323 citations
- Open data quality measurement framework: Definition and application to Open Government Data · Government Information Quarterly · 2016 · 312 citations
- Schema.org · Communications of the ACM · 2016 · 306 citations
- Data Derivatives · Theory Culture & Society · 2011 · 303 citations
- Data science empowering the public: Data-driven dashboards for transparent and accountable decision-making in smart cities · Government Information Quarterly · 2018 · 278 citations
Most recent work
- Data Architecture Maturity as A Predictor of Enterprise AI Success in Regulated Industries · International Journal of Advanced Artificial Intelligence Research · 2026
- Connecting the herd and the habitat: A plea for an AI-driven framework for integrating livestock and biodiversity data · One Ecosystem · 2026
- A visual analytic method to investigate patterns and structures of missing data · PLoS ONE · 2026
- Cleansera: A Context-Aware, Algorithm-Centric Data Cleaning System with RAG-Enhanced Intelligence · International Journal for Research in Applied Science and Engineering Technology · 2026
- An Artificial Intelligence-Based Data Integration Framework for Real-Time Cross-Source Data Harmonization · Journal of Computing and Data Technology · 2026
- Data Mapping Framework for the AAS · atp magazin · 2026
- Auto DW: An Agentic LLM-Based System for Automated Data Wrangling and Excel Intelligence · Zenodo (CERN European Organization for Nuclear Research) · 2026
- Comprehensive insights into bitemporal databases: a PRISMA-guided systematic literature review · Journal of Data Information and Management · 2026
- Generalized Entity Matching with Adaptivity via Large Language Models · Proceedings of the ACM on Management of Data · 2026
- A database-driven research data framework for integrating and processing high-dimensional geoscientific data · Geoscientific Instrumentation, Methods and Data Systems · 2026
Find a gap in your own Data Quality and Management sub-topic
This page shows what the Data Quality and Management literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →