Privacy-Enhancing Technologies Let AI Use Real Data Safely
Synthetic data can't replace real-world information for high-stakes decisions, but PETs enable secure collaboration without exposing sensitive records.

The real-data dilemma in AI development
Artificial intelligence systems deliver their greatest value when trained on recent, accurate, and representative data. Yet that same data—financial records, healthcare histories, behavioral patterns—is also the most sensitive and tightly regulated. Organizations face a fundamental tension: the information needed to build effective AI models is precisely the information they cannot freely share.
According to Chris Botha, CEO and Co-Founder of Omnisient, writing for the World Economic Forum, this challenge is particularly acute in finance, healthcare, and fraud prevention, where fragmented datasets across different institutions each capture only a partial view of individuals. A bank understands repayment behavior, a retailer sees household spending, and a telecommunications provider tracks digital service access. The predictive power lies in how these signals relate to one another, but combining them safely remains difficult.
Why synthetic data falls short
Synthetic data—artificially generated information designed to mirror real-world patterns—has emerged as one proposed solution. It serves legitimate purposes in testing systems and exploring edge cases, but cannot carry the full weight of high-stakes AI applications.
The fundamental limitation is that synthetic data only reproduces what is already known. It cannot capture emerging behaviors, unexpected correlations, or subtle characteristics present in live data but absent from the assumptions used to generate it. Where real data is sparse, synthetic versions inherit the same gaps and biases without making them visible. Germany's Federal Ministry of Finance reportedly proposed allowing tax authorities to use real taxpayer data for AI development after determining that fictitious data proved ineffective, though the proposal raised privacy concerns.
Even models trained entirely on synthetic data must ultimately assess real applicants, transactions, or patients at the point of decision.
How privacy-enhancing technologies work
Privacy-enhancing technologies offer a different approach. Rather than replacing real data with synthetic alternatives or copying sensitive records between organizations, PETs enable analysis without exposing or exchanging the underlying information.
The technical implementation uses tokenized identifiers to match records between organizations, allowing them to establish shared customers without revealing identities. Analysis occurs inside neutral environments where raw records remain invisible to all parties. Only aggregate results, model features, or individual scores leave the secure environment.
A practical demonstration in South Africa illustrates the potential. The continent's largest grocery retailer and leading banks tested whether shopping behavior could predict creditworthiness without sharing consumer data. Using anonymized and tokenized datasets matched in a privacy-preserving environment, they scored eight million previously credit-invisible people. The analysis qualified 3.2 million individuals for affordable credit they would otherwise have been declined, with models using grocery data showing a 41% accuracy improvement.
This outcome was only possible because financial institutions could access anonymized grocery data through PETs. The predictive signals had never been observed by the banks making lending decisions, and no synthetic process could have produced them.
Why it matters
The obstacle to better AI is no longer purely technical—it's about accessing real data while respecting legitimate privacy constraints. PETs don't eliminate the need for consent or regulation, but they address the central barrier to collaboration on sensitive information. As adoption barriers like cost, complexity, and absent standards begin to recede, these technologies expand what data can be used in exactly the areas where models are weakest: thin-file consumers, emerging fraud patterns, and underserved patient populations. The path to more accurate AI runs through real data, not around it.
These details were first reported by the World Economic Forum.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call
