ETSI Publishes 18 Metrics to Measure AI Data Quality
New technical standard defines formulas to assess completeness, fairness, and privacy before datasets enter production systems.

The European Telecommunications Standards Institute has released TR 104 180, a technical report establishing 18 standardized metrics that organizations can use to evaluate whether their datasets meet quality thresholds before deploying them in AI systems.
The standard provides specific formulas for calculating each metric, giving teams a quantitative method to assess data fitness rather than relying on subjective judgment.
Four categories of measurement
ETSI organizes the metrics into four distinct groups. The first covers foundational attributes: completeness, accuracy, consistency, and the absence of duplicate records. The second group addresses usability factors, including availability when required, sufficient documentation for provenance tracking, and currency of information.
The third category focuses on fairness, examining whether datasets represent different demographic groups equitably. Privacy considerations form the fourth group, evaluating both the risk of individual identification and the protection of sensitive attributes.
"It is essential that data quality is measurable, especially for organisations who need to establish whether its data is fit to essential intents, like it would be the case of trustworthy AI," said Diego Lopez, Chair of the ETSI Technical Committee DATA.
Testing reveals census dataset flaws
Researchers validated the metrics against two public datasets. Aircraft engine sensor data performed well across all measures, demonstrating completeness, accuracy, and temporal stability.
A widely-used US census dataset designed to predict income levels above $50,000 revealed significant problems. The data showed approximately 31 percent of men classified as high earners compared with roughly 11 percent of women—a nearly threefold disparity that TR 104 180 flags as a bias indicator.
The same census dataset failed two distinct privacy tests. First, combining just four fields—age, race, sex, and country—proved sufficient to uniquely identify specific individuals within the dataset. Second, sensitive personal information appeared in plain text without masking or encryption.
Why it matters
As regulators worldwide impose requirements for AI transparency and fairness, organizations need objective methods to demonstrate data quality. TR 104 180 treats privacy and fairness as measurable quality dimensions alongside traditional technical metrics, reflecting the reality that biased or privacy-compromising datasets create business and legal risk regardless of their technical accuracy. The standard provides a common framework that procurement teams, auditors, and data scientists can reference when evaluating datasets for production use.
Open-source implementation available
The working group behind TR 104 180—which included Sejong University, EGM, TTA, Daejeon University, and CNIT—developed an open-source tool that scores datasets against all 18 metrics. This implementation allows organizations to apply the standard without building measurement infrastructure from scratch.
The details were first reported by Help Net Security.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call