Last updated on September 18th, 2026 at 04:07 pm
Suppose you work in a hospital. Suppose you work at a hospital. Data that has been collected over the years, which you may have, that you might be able to use to train a machine learning model that could detect cancer earlier than any current diagnostic test. However, it is not possible to deliver these data without compromising Patient privacy, GDPR, and patient trust – not with a Research Lab, not even with a cloud server.
This is not a question of IF – this is a question of WHEN. Hospitals, banks, and governments are facing it, and so are they. That’s why Privacy-Enhancing Technologies (PETs) have moved from papers to real systems.
This isn’t a cryptography tutorial on encrypting simple text messages. A first-floor view with an honest, realistic account of what PETs can actually achieve now, where they still fall short, and what will be delivered in the next 2-3 years. If you’re any of the above, or simply someone with careful thoughts on how AI systems are constructed, this breakdown will be useful to you all.
Table of Contents
Privacy-Enhancing Technologies (PETs): Beyond Encryption – The Mental Model You Actually Need
For most, the term privacy tech evokes encrypted databases, strong passwords, and perhaps a VPN. But PETs are a different category.
They’re less about preserving data when it’s not moving from one location to another. They’re about protecting data when it’s actively used, on, run on, queried, trained with. That’s where classical encryption cannot be used.
A helpful three-layer way to think about it, drawn from the Privacy-Enhancing Technologies 101 frameworks published by OECD and industry bodies:
- Ensure that data are protected at rest / in transit: Standard encryption / TLS / key management – basically addressed.
- Protect data in use: PETs at work (Federated Learning (FL), Differential Privacy (DP), Secure MPC, Homomorphic Encryption (HE), Secure Trusted Execution Environments (TEE))
- Allow for controlled release – DP-protected reports, synthetic data, aggregation-only APIs – useful outputs that do not share out individual records.
PETs are not a single technology. They’re a family of methods, each appropriate to different threat models, data types, and performance requirements. The error teams can make is using them interchangeably.
What’s Already Running in Production (2025–2026 Reality Check)
A lot of discussion on “PETs – the future of privacy.A lot of chatter around ‘PETs – the future of privacy’. A couple of them, however, are already in use, under the radar, in systems that you likely use every day.
Differential Privacy – Accuracy vs Privacy, Tuned Daily
Requirements compromise on accuracy, and is their accuracy fair in terms of privacy?
If you’ve ever interacted with the Google telemetry report or Apple’s iOS usage analytics, you have met Differential Privacy. In DP, noise is mathematically introduced into the data output so no single individual’s data record can be gleaned from the output, even by someone with access to external data sets.
National statistical offices use it for census releases. It is also available for companies in the tech sector for behavioral analytics. The (epsilon) ε parameter in the core knob regulates privacy strength and the level of noise in the voting result. Lower valε values mean stronger privacy and more noise in the voting result. I have seen teams set an epsilon that, for some notion of the DP definition, they would adhere to. Still, the resulting statistics were so degraded they were operationally useless. Tuning matters enormously.
DP is mature. Not perfect (with repeated queries and misconfigured pipelines there are real risks!), but is the most production-ready PET for analytics and telemetry outputs.
Federated Learning – Training Without Moving Data
Federated Learning is the idea of teaching your keyboard to guess your phrase without Google even seeing your every keystroke. The model trains are local on your device, all gradient updates are not – but only – data changes (not models) leave your device. Increasingly, FL & MPC are being combined in healthcare and financial pilots.
FL is used in mobile keyboards and recommendation systems at hospital networks, as well as in healthcare research pilot initiatives. But most introductory articles miss one crucial point: FL doesn’t always ensure privacy. Inference attacks can leak training data through gradients. The standard production fix is to recode the system with both FL and DP (clipping and noise), then capture the production with secure aggregation to ensure the server sees summed data, not individual updates.
My experience showed that even well-designed FL systems can be susceptible to poisoning attacks if participating nodes are not properly authenticated. The structure is good; the deployment details are important.
Trusted Execution Environments – Hardware as the Last Line of Trust
Today, all the major cloud vendors (AWS, Azure, Google Cloud) support confidential computing instances from trusted execution environments (TEEs): Intel SGX, AMD SEV, ARM TrustZone. They are hardware-created enclaves and part of a memory region inaccessible to the cloud operator.
This is used in sensitive applications such as AI inference, medical data analysis, and financial fraud detection, where workloads run on third-party infrastructure. Still, you can’t fully trust the infrastructure provider.
The downside: TEEs have been targets of side-channel attacks, such as Specter, Meltdown, and more specific enclave attacks like the AEPIC Leak. You need to watch vendor patches closely. Unlike a “set and forget” device, TEEs are a powerful control and one that must be maintained to mitigate threats.
Synthetic Data and Secure Aggregation — Already Mainstream
Synthetic data generation has become popular for sharing external test data, training models when real data cannot leave a jurisdiction, and regulatory compliance testing.
More limited types of MPCs have been used for secure aggregation in privacy-preserving ad measurement (Apple’s Private Click Measurement, Google’s Privacy Sandbox), cross-bank fraud prevention, and federated analytics platforms.
Quick Reference: PET Landscape at a Glance
| Differential Privacy | Telemetry, analytics releases | Accuracy trade-off with low ε | Production |
| Federated Learning | On-device / silo training | Gradient leakage risk | Production |
| Homomorphic Encryption | Encrypted computation | High compute cost | Emerging |
| Secure MPC | Multi-party analytics | Communication overhead | Narrow production |
| TEEs / Confidential Compute | Cloud-sensitive workloads | Side-channel attacks | Production |
| Synthetic Data | External data sharing | Can preserve bias | Production |
What’s Just Beginning – The Next 2–3 Years
At this point, things get really interesting and really unpredictable.
General-Purpose Homomorphic Encryption — The Holy Grail Getting Closer
Randy, thanks to you, I can see how to improve general-purpose homomorphic public key encryption. Thanks to you, I can see ways to improve general-purpose homomorphic public key cryptography, the Holy Grail.
The PET with an esoteric name: Homomorphic Encryption: arbitrary computations can be performed directly on the encrypted data, and the result is the same as if the computations were performed on plaintext. No one involved in the computation knows any real numbers. If you dig into the details of Homomorphic Encryption, it may be clear why it is still largely a research-to-production story in 2025.
The difficulty: HE is expensive, given the computer. Milliseconds of operations can become seconds or minutes on ciphertext. Today, deployments are more practical: restricted static encryption of operations in a specific pipeline instead of workloads as a whole, mixed with TEEs and MPC to hide the costly HE operations in the most sensitive part of the pipeline.
The top 3 open-source HE libraries are Microsoft SEAL, OpenFHE, and TFHE-rs. General-purpose HE for real-time AI inference is a few years away, but with hardware acceleration, Mario is improving year to year.
Composed PET Stacks — FL + DP + TEE + Synthetic Data Working Together
This new architecture pattern is not “pick one PET”—it’s stacking several PETs to suit the threat model. A healthcare AI pipeline could involve data remaining within healthcare silos (FL architecture), DP-clipping gradients before they exit each silo, aggregating within a TEE on a cloud server, and validating output model results with synthetic data before releasing them beyond the healthcare environment.
Such a deliberate arrangement is currently being structured in Health Care networks, financial fraud rings, and cross-border data analytics initiatives. The engineering difficulty is complex and heavy — the complexity of each PET, the attestation requirements of each PET, and the debugging load of each PET. The governance story is much purer, however – you can show at each layer what privacy guarantee it affords.
PETs for IoT and Edge Computing
However, IoT brings constraints that traditional PET frameworks don’t address, such as intermittent connectivity, limited/embedded compute resources on edge devices, and mobility across network topologies. Papers for 2024-2025 are thus benchmarking FL frameworks and DP implementations on the edge.
This is an area to watch! The number of Lightweight PETs for healthcare wearables, industrial sensors, and smart cities will rise in the coming years, requiring significant development.
PETs as Data Sovereignty Infrastructure
One of the least discussed applications of PETs is data sovereignty, where organizations can use global cloud analytics without losing control of data that cannot leave a jurisdiction, with cryptographic guarantees helping ensure compliance. For instance, EU-based business entities with data processed through US cloud-based infrastructure are increasingly adopting TEEs and MPC to demonstrate that data remains covered by EU law when processed outside its borders.
The Real Challenges – Not Just Technical Ones
All the PET discussions seem to focus on cryptography, as I saw during mapping of real problems. The tougher problems are generally organizational.
The Utility-Privacy Trade-off Is Fundamental, Not a Bug
Increased privacy protection generally comes at a cost of decreased accuracy or decreased usefulness of the data. The downside of DP with a very low epsilon is that the results are often assumptions with high certainty, which can lead to a noisy data set and inaccurate conclusions. FL models trained from miniaturized disconnected silo data sets can perform much worse than centrally trained models.
There is no fat grips on the evening train. Calibrating “how much privacy is enough?” is not a technical question, but one that lawyers, organizations, and ethicists could answer. Sometimes, teams misunderstand epsilon selection and approach it only from a technical perspective.
Attacks on PETs Themselves
- Gradient inversion attacks can recover training data from FL gradient updates, particularly with large batch sizes.
- In an inversion attack, it is possible to infer information about the training data from the model for use in prediction.
- Membership inference attacks are used to determine if a certain record existed in a model’s training set.
- Synthetic data generators can learn about an individual or a small group of individuals in the source data and leak those individuals into the synthetic generation.
However, if DP isn’t set up properly, multiple DP mechanisms can lead to privacy budgets far lower than desired.
Every PET will have an attack surface. Without a threat model and adversarial evaluation, a PET deployment is not a step up in privacy; it’s security theater.
The Skills Gap Is Severe
Barriers to Adopting PETs is a new study by the U.K. Information Commissioner’s Office (ICO) that confirms what many practitioners know: it’s hard to find individuals proficient in both cryptography and machine learning, and they need that expertise to deploy PETs properly. Most ML engineers don’t have cryptography expertise. Most cryptographers lack ML expertise. The field now needs individuals who can connect both.
This is also a content gap: practical, concrete knowledge about deploying PETs is lacking beyond academic abstraction. Most implementation guides do not include actual code or have a “real threat model breakdown.
Regulatory Uncertainty – When Is DP-Protected Data Still Personal?
One of the most tantalizing unanswered questions is: is data processed using DP, when collected before processing, still considered “personal data” under GDPR? But regulators say the answer is: it depends. However, residual re-identification risk, particularly if DP outputs use linked datasets from external sources, might result in data still being legally considered personal. At present, the area of PETs for Regulatory Compliance is a work in progress, as law and technical aspects are still under development.
PDEs must hire privacy attorneys at the outset, not at the end of the process, when deploying PETs.
How Practitioners Can Actually Use This Right Now
Map PETs to Threat Models, Not Buzzwords
The question is not “which PET should I use?” But what is my real threat model? If there are multiple responses, then there will be multiple PETs:
- Does the data need to remain physically located somewhere? → Federated Learning.
- Want an extra layer of privacy for public statistics or model output? → Differential Privacy.
- Need to do computations on untrusted cloud providers? → TEEs.
- Collectively compute both without revealing any information about the input to others? → Secure MPC or Homomorphic Encryption.
- Have to share a dataset externally for testing or pre-training? → Use Synthetic data, generated without the risk of identification.
Start Narrow and Prove It Out
The organizations I have observed that have been successful with PETs began with a single, high-value pilot with one clear focus. Two hospitals experimented with federated learning of a specific diagnostic model. Next, a financial consortium is implementing secure aggregation for a narrow fraud signal—a single analytics team, with one telemetry pipeline deployed using DP.
Supporting a narrow scope provides you with a real system to learn from, a concrete privacy analysis to inform regulators, and governance patterns to reuse. Building a full PET stack on a single data platform is often a huge task that is hard to sustain.
Open-Source Frameworks Worth Knowing
- Google’s production-tested FL framework, TensorFlow Federated (TFF), with support for DP and secure aggregation.
- OpenMined’s privacy-preserving ML framework, PySyft, supports FL, DP, and MPC.
- Microsoft SEAL / TenSEAL — HE libraries with Python bindings to experiment with encrypted computation.
- Relevant library: OpenDP — open-source DP library used for the applications of the US Census at Harvard.
- Flower (flwr) — Framework-agnostic FL library for PyTorch, TensorFlow, and JAX.
- All of these are available for free and actively developed. If you don’t have small but real examples to build and you watch, you won’t build real credibility in this space.
Frequently Asked Questions
Do PETs replace security and compliance frameworks like GDPR?
No. PETs are controls that complement legal compliance, IAM, logging, and incident response. So OECD and regulators can see that PETs aren’t a compliance quick fix. They diminish privacy risk – they do not remove legal obligations.
If I apply Differential Privacy, is my data automatically anonymous?
Not necessarily. If the information remains identifiable in any way after DP has been applied, it will still be regarded as “personal” information for GDPR – even if it is anonymized even if the personal information remains identifiable in certain ways after DP (such as through linking with external datasets), it will still be considered “personal” for GDPR purposes. Regulators also show growing skepticism about ‘naive’ anonymization claims. For privacy, an appropriate (and well-selected) epsilon can often make a massive difference in lowering the risk of re-identification, and compositional considerations play a crucial part, but this does not require toggling any privacy flag.
Which PET is best for training AI on sensitive data?
It depends on the constraint that it operates under. Where data must remain on-site: FL with DP is the standard production approach, with secure aggregation. TEEs are the pragmatic solution if the computation must run on untrusted cloud compute. When several independent parties need to compute together but neither party can access the other’s data, MPC or HE can be used. The correct answer is usually a mixture of two or three of these.
Are PETs too slow for production?
Some are in their general-use form. However, general HE and rich MPC protocols can still be too slow for many real-time applications. Many workloads can be production viable for DP, FL, and TEE-based confidential computing, provided some care is taken to make sure operations, especially memory operations, have low latency and that large PETs are scoped to the most important stages of the pipeline.
What’s the most practical starting point for a team new to PETs?
Choose one specific and easy-to-understand use case with a clearly identifiable threat model. Initially run a small FL + DP pilot or deploy a single, sensitive compute workload that requires confidential computing. Have privacy experts on board from the beginning. List the conditions or guarantees your selected PET will provide and not provide for an overview of compliance and legal issues. Then, develop reusable governance structures.
My Take: What This Space Actually Needs
PETs are past the research-only stage. The foundational features of DP, FL, TEE, and secure aggregation are in production at scale. Not the encryption; it’s deployment ignorance, risk-model awareness, and compliance confusion.
The missing element is practitioners who can balance attack exposure, governance needs, and practical realities to provide a thorough understanding of a PET architecture without resorting to business management or theory.
If you are developing some real depth here, you are either an 18- to 35-year-old building in AI, or perhaps you’re a data engineer, or perhaps you’re seriously paying attention to just how it ought to be performed within the era of large-scale device learning. The tools are free and open source. The study is open source. There is a significant need for those with technical and governance skills and an increasing need for those who have both.
Organizations with the largest budgets are not the ones succeeding in bringing PETs to fruition. These are the people who started with a very narrow view, built real systems, and brought a legal team and a technical side into the same process from the start.
That’s the actual playbook. All the noise is worthless, except for that.
Read:
Kubernetes Security – Hardening Cloud-Native Workloads
I’m a technology writer passionate about AI and digital marketing. I create engaging and useful content that bridges the gap between complex technology concepts and digital technologies. My writing makes the process easy and engaging. I encourage participation I continue to research innovation and technology. Let’s connect and talk technology!



