Data Quality: the Key to Trustworthy AI in Care
AI holds great promise for healthcare - but trust must be earned. One key ingredient is often overlooked: data quality. Without high-quality data, there can be no reliable AI. And that’s precisely where things often go wrong, because clinical data is typically gathered in fast-paced, complex environments, often under pressure.
To make AI truly work in healthcare, we need to prioritise quality from the very beginning. That means involving everyone who works with data, says Jens Declerck of the European Institute for Innovation through Health Data (i-HD), and also: embracing a lifecycle approach, and ensuring representativeness. Only then can AI live up to its promise: better and fair healthcare for all.
The importance of data quality - and what happens without
AI learns from the data we feed it. But if that data is incomplete or unevenly distributed, AI can end up amplifying the very inequalities it’s meant to reduce.
Studies have shown time and again: bias in AI often stems from poor data quality. Common issues include:
-
Missing or incomplete information
-
Inconsistent data recording practices
-
Underrepresentation of certain population groups
As long as we ignore the foundations of healthcare data quality, we risk building AI systems that embed, rather than challenge, existing biases.
Bias in hospital readmission prediction
A 2023 study evaluated the fairness of popular models for predicting 30-day hospital readmission, including the LACE index and HOSPITAL score. The findings were clear: these models consistently underperformed for black patients and people with lower incomes. Despite having similar clinical profiles, these patients were less likely to be flagged as a readmission high-risk. The problem wasn’t just the algorithm. It was the data. Social health factors, such as housing instability or lack of follow-up care, were often missing or only partially recorded in the patient files. Moreover, for marginalized groups, data was frequently incomplete and inconsistently coded. Instead of identifying true medical risk, the models learned to mirror existing inequalities.
EU AI Act - Art. 10: Data & Data Governance
Any AI system that is considered high-risk must, under the European AI Act (Article 10), be trained, validated and tested using high-quality data. The EU sets strict requirements to ensure that the data is reliable, representative and fair:
-
the data must be relevant, complete, sufficiently representative, as error-free as possible, and aligned with the purpose of the AI system being developed
-
the data must be properly managed, with clear procedures for collecting and storing the information, processing it (such as labelling, cleaning...), and applying techniques to detect and address biases or missing data
-
the data must take into account the specific geographical, contextual, functional and behavioural setting in which the system will be used
-
if strictly necessary to prevent discrimination, sensitive personal data (such as ethnicity, religion, health...) may be used exceptionally, but only under strict conditions and with robust safeguards
Data Quality needs a Lifecycle Approach
Improving data quality isn’t a one-off task. It’s a continuous effort that involves everyone: from frontline care providers to AI developers. Frameworks like METRIC can help assess datasets for completeness, consistency, and representativeness, but they are rarely used in day-to-day clinical practice.
A robust, lifecycle-based approach checks data quality at every stage:
-
Data entry: use structured fields, required inputs, and alerts for inconsistencies
-
Data processing: apply validation rules, track provenance, conduct regular audits
-
Modeling: test for demographic representation and context-aware outcomes
-
Monitoring: implement dashboards that flag data drift or missing values in real time
A 2024 review found that while tools exist, few healthcare institutions use them in an integrated way. Quality checks remain fragmented and static, and the infrastructure for dynamic, end-to-end monitoring is still lacking.
Despite the urgent need to turn theory into practice, adoption remains low. Most clinical environments have yet to move beyond academic or policy-level engagement with data quality.
How to Embed Data Quality across the Lifecycle
To make AI systems fair and effective, healthcare organisations should:
- Implement quality checks from the very first data point
- Support clinical teams with training and user-friendly tools
- Align management incentives with data stewardship
- Complement frameworks like METRIC with continuous monitoring and contextual insight
The Importance of Representative Data
An often underestimated, but crucial aspect of data quality is the representativeness of datasets. When this is overlooked, models risk underperforming or discriminating once deployed in real-world settings.
Ensuring representativeness requires careful attention throughout the AI lifecycle:
-
Clearly define the target population and clinical setting
-
Monitor how cleaning or merging of data impacts demographic balance
-
Understand the clinical context in which data is recorded, as this determines which data is collected, how consistently it is documented, and for what purpose. This includes:
-
Clinical workflows
-
Documentation practices
-
Organisational priorities
-
Without this context, AI may misinterpret signals and generate outputs that are misaligned with reality.
Monitoring representativeness
Methods for achieving representativeness as a routine process, rather than a one-off checklist, are still evolving.
The PARADISE project, for example, aims to improve the transparency of the training data used in AI systems. One of its key tools is a "data quality fingerprint": a structured method to document and assess the composition and quality of AI training datasets, with a strong focus on representativeness.
Trustworthy AI Is a Shared Responsibility
Building trustworthy AI is about more than smart algorithms. It’s about making thoughtful, responsible choices where data is concerned: from the registration on the healthcare floor to the training and (continuous) monitoring of the model.
This is a collective effort. Care providers, developers, policymakers, and data managers must collaborate to construct AI systems that are grounded in quality. With a shared vision, practical tools, and lifecycle awareness, we can develop AI that truly support healtcare, for everyone.
References
- Wang HE, Weiner JP, Saria S, Kharrazi H. Evaluating Algorithmic Bias in 30-Day Hospital Readmission Models: Retrospective Analysis. J Med Internet Res. 2024 Apr 18
- Hasanzadeh, F., Josephson, C.B., Waters, G. et al. Bias recognition and mitigation strategies in artificial intelligence healthcare applications. npj Digit. Med. 8, 154 (2025).
- Feng Chen, Liqin Wang, Julie Hong, Jiaqi Jiang, Li Zhou, Unmasking bias in artificial intelligence: a systematic review of bias detection and mitigation strategies in electronic health record-based models, Journal of the American Medical Informatics Association, Volume 31, Issue 5, May 2024, Pages 1172–1183
- Schwabe, D., Becker, K., Seyferth, M. et al. The METRIC-framework for assessing data quality for trustworthy AI in medicine: a systematic review. npj Digit. Med. 7, 203 (2024).
- Declerck J, Kalra D, Vander Stichele R, Coorevits P. Frameworks, Dimensions, Definitions of Aspects, and Assessment Methods for the Appraisal of Quality of Health Data for Secondary Use: Comprehensive Overview of Reviews. JMIR Med Inform. 2024 Mar 6
- Declerck J, Vandenberk B, Deschepper M, Colpaert K, Cool L, Goemaere J, Bové M, Staelens F, De Meester K, Verbeke E, Smits E, De Decker C, Van Der Vekens N, Pauwels E, Vander Stichele R, Kalra D, Coorevits P. Building a Foundation for High-Quality Health Data: Multihospital Case Study in Belgium. JMIR Med Inform. 2024 Dec 20
Read / learn more
How can AI be used in healthcare?
AI is rapidly transforming the healthcare domain. From diagnosis to treatment: AI helps doctors and other medical professions to take better-informed decisions and to supply more personalised care.
How to make AI reliable and transparent
A transparent AI system creates support, strengthens trust and keeps risks manageable. Eulaly Vanroelen from Thomas More shows you the way to a fair AI system in 7 practical steps.
AI in Healthcare. Hype or Help?
course - online - KU Leuven, VAIA
From enigma to insight: How explainable AI can help us unravel biological complexity
In this blog by BioLizard, we explore how explainable AI can shed a light on AI's decision-making process in computational biology.
Jens Declerck
Expertise:
- Data quality and biases towards AI
- Explainability using data quality assessments
- Understanding data sources towards extracting and harmonizing data for training AI
- How AI might bring value in hospitals and specific towards creating more data quality maturity
Related courses
Fundamental concepts of multi-omics data integration
Opleiding - Leuven - Genomics Core Leuven
Transforming Healthcare with AI
Two-day program - Cambridge - MIT Sloan School, MIT Jameel Clinic
HPC Café
Seminarie - Online, Karlsruhe - KIT, HammerHAI, Training geselecteerd in de Belgische AI Factory Antenna (AIFA)