Why there is no single percentage
95% valid data can be excellent for a newsletter and not enough for invoicing or training an AI system. The threshold depends on the risk of using bad data, not on a general reference figure.
How to set a threshold in three steps
- Identify the use: which decision or process depends on the data.
- Estimate the impact of an error: money, time, customers affected, legal risk.
- Choose a minimum level per dimension and write it down together with who reviews it.
Examples of criteria
| Use of the data | Impact of an error | Requirement |
|---|---|---|
| Invoicing or payments | High: money and claims | Very high |
| Personal data (GDPR) | High: rights and fines | Very high |
| Management reports | Medium: wrong decisions | High |
| Informational marketing | Low: minor annoyance | Moderate |
Review your thresholds
A threshold is not forever. If you meet it effortlessly for months, you may be able to tighten it. If you never reach it, either it is unrealistic or there is a root-cause problem to fix with data cleansing.
If you work with high-risk AI, Article 10 of the AI Act requires relevant and representative data, but it does not set a percentage: you must justify your own thresholds.
Quick answers
What is a good data quality percentage?
There is no universal percentage. It depends on the risk of use: data affecting money or people demands very high levels, while informational data tolerates less.
Who decides the acceptable level?
Whoever uses the data to make decisions, supported by the data owner. It should not be set by the technical team alone.
Does the AI Act set a quality percentage?
It sets no figure. It requires data for high-risk systems to be relevant and representative, and that you can justify your criteria.