Decision Sciences · Research topic

Open research questions in Data Quality and Management

33 unresolved questions extracted from the limitations and future-work sections of 437 Data Quality and Management papers in our library. Each links back to the study that raised it.

What the literature leaves open

  • Existing frameworks treat data quality dimensions as context-independent and universally applicable, yet no systematic method exists for surfacing and interrogating the latent value commitments and normative assumptions embedded in quality assessment frameworks themselves. Prior work identifies multiple dimensions (accuracy, completeness, etc.) but does not examine how the selection and weighting of these dimensions encodes particular conceptual frameworks and priorities that may systematically distort phenomena when applied across different use contexts.

    The DaTUM framework: a multi-sector thematic analysis of data quality dimensions and their impacting factors · 2026 · DOI
  • Finally, the recommendations made for mitigating the impacting factors should be explored, such as undertaking studies that utilise machine learning algorithms to find and correct outliers or performing data stakeholder analysis to understand how to best communicate uncertainties and achieve project outcomes. A key limitation of the study is the imbalance of our sample towards the higher education and automotive sectors.

    The DaTUM framework: a multi-sector thematic analysis of data quality dimensions and their impacting factors · 2026 · DOI
  • directions Bitemporal model type Type of model used: RQ2 The methodology of this study is designed as a hybrid sys- tematic literature review approach. The first layer consists of bibliometric and scientometric analyses, which support the overall review rather than serving as independent methods (Hilmi et al. 2023). It employs VOSviewer and BibExcel as tools for bibliometric and scientometric mapping (Dou- lani 2021). The second layer's primary contribution lies in the qualitative synthesis of the selected literature, including a comparative evaluation of bitemporal models, an assess- ment of DBMS infrastructure support, analysis of domain- specific applications, and the identification of research gaps. The second layer applies the PRISMA framework and the GQM approach as methodological frameworks to guide the qualitative synthesis and systematic data extraction from the 54 included studies (Farina et al. 2022). Thus, integrating these two layers in a hybrid approach enables a more comprehensive understanding of the field by linking macro-level research patterns to micro-level tech- nical insights. Such combined approaches are increasingly recognised in systematic literature review practices, where bibliometric techniques are employed to strengthen, rather than replace, structured qualitative synthesis (Ding and Meng 2014b).

    Comprehensive insights into bitemporal databases: a PRISMA-guided systematic literature review · 2026 · DOI
  • The paper states future work will evaluate Cleansera's performance and empirical efficacy across industry datasets, but does not specify which industry types (healthcare, finance, e-commerce, manufacturing), dataset sizes, data quality profiles, or cleaning rule complexities will be tested. Comparative performance evaluation against existing data cleaning systems on standardized benchmarks is not outlined.

    Cleansera: A Context-Aware, Algorithm-Centric Data Cleaning System with RAG-Enhanced Intelligence · 2026 · DOI
  • The flowchart-driven execution model is proposed as a design principle for deterministic AI-assisted decision-making, but no formal verification methodology, state machine specification, or branch coverage analysis is provided. Testing of error handling and recovery paths across all conditional branches in authentication, context detection, and cleaning workflows requires specification of failure scenarios and system behavior.

    Cleansera: A Context-Aware, Algorithm-Centric Data Cleaning System with RAG-Enhanced Intelligence · 2026 · DOI
  • The RAG-enhanced intelligence component mentioned in the paper title is not detailed in the provided excerpt, and no specific evaluation metrics, retrieval-augmented generation architectures, or knowledge base configurations are described. The integration of RAG with the context detection and cleaning algorithms requires empirical comparison against non-RAG baselines on domain-specific datasets.

    Cleansera: A Context-Aware, Algorithm-Centric Data Cleaning System with RAG-Enhanced Intelligence · 2026 · DOI
  • Cleansera's cleaning pipeline includes seven sequential stages (Schema Validation, Duplicate Detection, Missing Value Treatment, Format Standardization, Outlier Handling, Semantic Validation, Quality Checkpoint), but the paper provides no empirical performance benchmarks or scalability analysis on datasets varying in size (from thousands to millions of records) or complexity (number of columns, data types, missing value percentages).

    Cleansera: A Context-Aware, Algorithm-Centric Data Cleaning System with RAG-Enhanced Intelligence · 2026 · DOI
  • The master field identification algorithm ranks fields using uniqueness score, semantic relevance, referential integrity, and type consistency, but the paper does not specify how these four ranking criteria are weighted relative to each other or how the ranking function performs on datasets with high cardinality columns, composite keys, or sparse referential integrity patterns.

    Cleansera: A Context-Aware, Algorithm-Centric Data Cleaning System with RAG-Enhanced Intelligence · 2026 · DOI
  • The data loss detection algorithm compares raw and cleansed data to quantify deletion rates and modification rates at record and attribute levels, but no specific metrics are provided for acceptable loss thresholds or how these rates should be interpreted across different data types and domain-specific cleaning rules. Validation of this dual checkpoint quality assurance process on heterogeneous industry datasets is explicitly identified as future work.

    Cleansera: A Context-Aware, Algorithm-Centric Data Cleaning System with RAG-Enhanced Intelligence · 2026 · DOI
  • The context detection algorithm assigns weighted confidence scores based on semantic matches and statistical indicators, but the paper does not specify the threshold values, weighting formulas, or how confidence scores are calibrated across different industry domains. Empirical validation of context detection accuracy rates across specific industry datasets is needed to determine when manual override is triggered.

    Cleansera: A Context-Aware, Algorithm-Centric Data Cleaning System with RAG-Enhanced Intelligence · 2026 · DOI
  • This paper argues four important HIB theories are insufficient for describing users' search strategies for data because of assumptions about the attributes of objects that users seek.

    Data, not documents: Moving beyond theories of information‐seeking behavior to advance data discovery · 2024 · DOI
  • This article concludes with suggestions for an inclusive computing tool or environment, additional research on the treatment of missing data, and reasonable and flexible interpretations of the WWC standards.

    Computing Tools for Implementing Standards for Single-Case Designs · 2015 · DOI
  • The measurability of quality dimensions varies substantially across contexts and assessors, yet no framework exists for understanding how conceptual adequacy—the degree to which a dataset's definitional scope and measurement framework align with the analytical question being asked—can be systematically evaluated. Current work shows that assessors disagree on dimensions like completeness but does not address how to evaluate whether the data's underlying conceptual framework is fit for a particular analytical purpose.

    On the measurability of information quality · 2010 · DOI
  • Because the quality-enforcing structures present in the MARC world–mature standards, common documentation, and bibliographic utilities–are lacking in the metadata world, metadata practitioners desiring to improve the quality of metadata used in their libraries must develop and proliferate their own processes of evaluation and transformation to support essential interoperability.

    Metadata Quality: From Evaluation to Augmentation · 2008 · DOI
  • Big data's exponential growth has inverted the traditional research logic where measurement frameworks are articulated before data collection; yet no method exists for retrospectively assessing whether large-scale datasets assembled without explicit conceptual frameworks carry systematic biases in their population boundaries, definitional scope, or measurement assumptions that render them inadequate for specific analytical purposes. Current quality frameworks assume data already exists and focus on improving known dimensions rather than surfacing latent conceptual inadequacies.

    Managing Organizational Data Resources · 2000 · DOI
  • Data reuse and repurposing across system boundaries is increasingly common, yet no framework exists for identifying and managing the systematic distortions introduced when datasets assembled for one analytical purpose—with embedded assumptions about population boundaries and definitional scope—are repurposed for different analytical ends. Current lifecycle models address quality engineering within defined contexts but do not address the relational renegotiation of quality required when inquiry assumptions change.

    Managing Organizational Data Resources · 2000 · DOI
  • Despite the growing availability of Arabic digital content, persistent issues such as linguistic ambiguity, inconsistent data standards, fragmented data ecosystems, and limited cross-sector integration continue to hinder effective data-driven decision-making.

    HADQEF: a hybrid AI framework for enhancing Arabic data quality and governance in alignment with Saudi Vision 2030 · 2026 · DOI
  • As the synthetic data field is still relatively new, norms and precedents are yet to fully emerge and develop, and standards and consistent approaches are lacking.

    Guiding principles for the provision of synthetic data · 2026 · DOI
  • Global livestock production represents a primary driver of planetary change, yet its full ecological and public health impacts remain poorly quantified potentially due to a deep-seated data divide.

    Connecting the herd and the habitat: A plea for an AI-driven framework for integrating livestock and biodiversity data · 2026 · DOI
  • Therefore, future studies are recommended to explore system scalability, enterprise platform integration, mobile-based operational monitoring, and more comprehensive evaluation methods such as usability testing, user satisfaction analysis, and long-term operational impact assessment.

    Design and Development of an Employee Database and Daily Reporting Information System Using Rapid Application Development (RAD) · 2026 · DOI
  • Finally, the study shows that the bitemporal approach addresses the key limitations of tradi- tional systems for handling historical data by storing both valid and transaction times.

    Comprehensive insights into bitemporal databases: a PRISMA-guided systematic literature review · 2026 · DOI
  • Consequently, the decision makers will be able to treat mixed data, numerical and categorical data, to explain and predict phenomena in the big data ecosystem.

    Decision-Making Enhancement in a Big Data Environment: Application of the K-Means Algorithm to Mixed Data · 2019 · DOI
  • One issue that may arise when using these decentralised administrative data is that categorical variables are underreported by some of the data suppliers, for instance to avoid administrative burden.

    Detecting Reporting Errors in Data from Decentralised Autonomous Administrations with an Application to Hospital Data · 2018 · DOI
  • This creates a void of options for a user-driven assessment of data quality when metadata are sparse or unavailable, as is often the case with citizen science and volunteered geographic information.

    Measuring Spatial Data Fitness-for-Use through Multiple Criteria Decision Making · 2018 · DOI
  • Thus, different from other bootstrapping methods that explore privileged, hard to obtain information such as self‐citations and personal information, our proposed method produces topnotch performance with no (manual) training data or parameterization and in the presence of scarce information.

    Self‐training author name disambiguation for information scarce scenarios · 2014 · DOI

Most-cited papers in Data Quality and Management

Most recent work

Find a gap in your own Data Quality and Management sub-topic

This page shows what the Data Quality and Management literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.

Open the Research Gap Finder →

Related topics in Decision Sciences

33 open questions have been extracted from the limitations and future-work passages of 437 Data Quality and Management papers in our library. Each one below links back to the study that raised it, so you can read the original claim in context.

Tools for your next paper

Compare the categoryHonest roundups of the AI research tools, ours listed alongside the alternatives.

Command palette

Jump anywhere, run any action.