Last updated on September 19th, 2026 at 08:45 am
And there is a space most cloud architects wouldn’t discuss publicly. You can encrypt your data at rest, lock down your APIs, and still have your model weights sitting in clear memory when inference starts. Any cloud administrator, compromised hypervisor, or insider threat has a window, and it is broader than most teams can comprehend.
That is the issue that Confidential AI is literally addressing. No privacy as such. Not compliance theater. True physically enforced isolation where even the infrastructure provider has no access to see what your model is doing with sensitive information.
This article breaks down architecture patterns already in production, those still being determined, and where things are headed. This is worth your time if you are building AI systems on cloud infrastructure and dealing with anything regulated or proprietary.
Table of Contents
What “Confidential” Actually Means in a Cloud AI Stack
When most people hear “confidential computing,” they think of encryption. That is part of it, but encryption alone isn’t enough for the computation itself.
Confidential AI specifically refers to the process of running your data pipelines, inference into models, fine-tuning jobs, and agents within Trusted Execution Environments (TEEs) – hardware-isolated regions where your model memory is encrypted, access to it is attested, and not even the host OS or the hypervisor can know what is happening in there.
Present primitives are:
- AMD SEV-SNP – hardware-level encryption of VM memory and attestation.
- Intel TDX – trust domain independent of VMs and hypervisor.
- Intel SGX – enclave-level isolation (smaller footprint, more restrictions)
- ARM Confidential Compute Architecture (CCA): ARM designed a mobile and edge-based confidential compute architecture (CCA).
- GPU encrypted and attested (NVIDIA Hopper/Blackwell/Vera Rubin GPU) — Enhanced encryption of GPU memory, with attestation, allowing secure inference at scale.
Now the key clouds have some confidential VMs on behalf of at least confidential VMs. Others, such as Google and Azure, have AI workload reference architectures. That’s the baseline.
The Four Patterns You’ll See in Every Reference Design
Single-Tenant Confidential Inference
This is the most common entry point. You bundle an existing LLM or model, place it in an encrypted OCI container image, and deploy it in a confidential VM. The model is decryptable only after the TEE passes attestation and the key management service (KMS) issues the decryption key.
It flows as shown:
Client Bottom-up Client to Attestation Service with a Bottom-up Architecture KMS with a Bottom-up Architecture Confidential VM (decorates image using model) with a Bottom-up Architecture Inference.
The model weights in clear memory are never exposed outside the TEE.
The confidential inferencing environment employed by Azure and the Red Hat OpenShift environment using containers sandboxed by the system can both be configured to use this pattern in some form. This direct mutual attestation step – both the customer and model provider attest to each other having this configuration – is what makes this substantially different than simply running in a VM.
Lesson point: This is where you are starting. Get attestation of encrypted model-image functionality running before you attempt anything more complex.
Confidential RAG and Fine-Tuning
One of the most obvious enterprise applications to this stack is Retrieval-Augmented Generation (RAG) using proprietary or regulated data.
The architecture works as follows: documents are preprocessed, embedded, and stored as vectors in a database that runs within a TEE or confidential VM. The orchestrator of the RAG – retrieval, prompt assembly, context injection – also executes in the same sheltered environment. The LLC endpoint may be both internal and external, preferably also attested.
To perform fine-tuning, the training loop itself can be implemented within the TEE. Gradients, updating of weights, the whole backprop cycle- none of it is known to the cloud infrastructure. This matters to financial institutions and medical organizations that prefer to fine-tune on internal corpora without exposing them to a cloud vendor.
Anjuna’s confidential computing whitepaper addresses this, with detailed examples of customer scenarios. What I found in my review of their architecture was that the decision of where to place the vector store- inside and outside the TEE- was one of the first spots where real implementation choices diverge from the clean implementation image.
Confidential Federated Learning and Clean Rooms
These include several parties: banks, hospitals, competing enterprises, etc. – all of them, as well as a TEE aggregator operating in the cloud, train locally and send encrypted updates up to a TEE aggregator operating in the cloud. The aggregator produces a world model. No raw data leaves any participant. The cloud provider cannot check the updates in clear.
A canonical example is Google’s financial fraud detection reference architecture. The TEE aggregation server is the linchpin. This trend makes cross-institutional AI collaboration both legally and technically possible in highly regulated sectors.
What complicates this even more than it may sound is attestation orchestration among participants, which introduces significant operational complexity. Each party must check the TEE configuration separately and submit changes.
Rack-Scale Confidential GPU Clusters
This is where things become truly new. NVIDIA’s Vera Rubin NVL72 provides a confidential security domain that unifies dozens of GPUs and CPUs connected via NVLink. The whole rack forms one TEE boundary.
In the past, GPU TEEs were memory-limited, which made training large models unfeasible. Rack-scale changes that. Full training and high-throughput inference may now be fully utilized within a GPU TEE, with near-native performance, and in a multi-tenant cloud environment.
While reading NVIDIA’s documentation, I noticed it still needs to determine how to guarantee these as multi-tenant isolation platforms fully. The hardware exists. The active development area lies above it: the policy and quota implementation layer.
Where the Kubernetes Layer Fits In
Both the Red Hat OpenShift sandboxed Container and the confidential-containers open-source project introduce TEE isolation into standard Kubernetes workflows. Pods run in a micro-VM supported by the TEE, with attestation-conscious scheduling and automatic key release.
This matters because it lets platform teams make TEE capabilities available to data science and ML engineers using standard K8S tooling. The TEE complexity is isolated below the platform layer. Engineers write normal workloads. The splendor of isolation occurs below them.
One of the clearest free resources on how this wiring works is the project’s architecture documentation on GitHub.
What’s Still Being Figured Out
The trends mentioned above are practical and can be implemented today. A few things are still moving, and it is important to be forthright about them if you’re making a decision that affects architecture.
Mutual and third-party scale attestation – The idea is simple: both the customer and the model provider attest to each other that their TEE and software stack conform to specifications. Both Red Hat and Anthropic have described this. Cross-cloud standardization is still emerging.
Access control in agentic AI: agentic AI agents make API calls, retrieve documents, trigger workflows, and assume the role of non-human identities. The Cloud Security Alliance (CSA) has identified this as a gap, although reference designs for platforms on which confidential agents operate are only just beginning to surface.
TEEs with DP, HE, and MPC – Each of the above solutions can complement the others. Combining them with TEEs to address gradient leakage and offer stronger statistical guarantees is also an active research direction that has yet to become a reference pattern.
Fragmentation of attestation – Intel, AMD, and ARM all have varying attestation formats. Cloud providers use their own attestation services at the top. Current multi-cloud (or hybrid) confidential AI designs require glue code, and no universal standard exists yet.
This is where the relation to Post-Quantum Cryptography Migration comes into play. The TEE attestation cryptographic primitives, including key exchange, signature verification, and certificate chains, are all classical.
Post-quantum algorithms will require migrating the attestation and key management layers of the confidential AI architecture to post-quantum algorithms as quantum threats to these primitives become reality. Migration isn’t a current project, but architects building confidential AI today should track it.
The Real Challenges Nobody Puts in the Headline
Memory and Performance Constraints
TEEs (SGX in particular) that operate only on CPU have strict memory constraints. Large models don’t fit. GPU TEEs can help, but they still add overhead because they combine encryption, differential privacy, and sophisticated multi-party networking. The gap between performance and non-confidential baselines has narrowed significantly, but not completely.
Observability Is Genuinely Hard
TEEs are transparent in nature. And that is it. However, that also means your customary logging, tracking, and incident-response tools don’t work the same way. Confidential workloads require SRE practices to be rethought afresh: what metadata can definitively and safely be emitted, how to structure logs to make them useful without being leaky, and how to debug with the enclave as a black box by definition.
This is what reference architectures continually underestimate, as I’ve seen when reviewing operational documentation from multiple vendors. The charts appear neat. The reality of the ops is sloppier.
Threats That TEEs Don’t Handle
TEEs prevent intrusion at the infrastructure level. They don’t protect against:
- Model inversion attacks – guessing training input based on model output.
- Membership inference – testing whether a particular record used to be in the training data.
- Timely injection – Control of model behavior by design of model inputs.
- Data poisoning – poisoning training data before it is transferred to the TEE.
These require ML-level defenses: differential privacy, access controls, rate limiting, input validation, and monitoring. Confidential VMs are not a substitute for these. Architects who assume running in a TEE provides complete security coverage are misjudging it.
Free Resources Worth Your Actual Time
Instead of enumerating all, here are the ones providing the highest number of architecture-level signals per hour:
- Confidential AI reference architectures of Google Cloud – the optimal starting point for federated learning and analytics patterns with diagrams.
- The AI security of confidential computing provided by NVIDIA with confidential computing documentation – a requirement when learning how to use the GPU TEE, Hopper/Blackwell/Rubin coverage.
- The secret Red Hat LLM inference guide is the most effective for image encryption and mutual attestation workflow examples.
- The whitepaper AI/ ML in Anjunas- concrete customer scenarios of RAG and fine-tuning architecture.
- CSA Data Security within AI Environments. – threat and control mapping across the AI lifecycle, including considerations of agent/NHI.
- The Cybersecurity of AI and Standardization of ENISA – a standards-oriented perspective that helps align compliance.
- GitHub confidential-data-containers architecture.md – the most open and free official technical documentation of Kubernetes-native TEE integration.
Between these, the full picture from 101 to architecture, plus challenge analysis, will be complete and not at all costly.
Who Should Actually Be Building This Now
Confidential AI isn’t futuristic; it’s already being used today. It can support specific applications today. The authorities that ought to be driving this now:
Healthcare and life sciences – educate about patient data, share models across institutions, train AI diagnostics with specific data residency criteria.
Financial services – fraud detection instrument for all institutions, proprietary model protection, regulatory data processing to inference workloads.
Legal and professional services – document analysis on confidential client data where the AI provider cannot be a de facto data processor.
Defense and government – classified workload processing, multi-agency data cooperation with severe need-to-know limitations.
When your workload includes proprietary model IP, regulated datasets, or data collaboration among multiple parties, the confidential AI stack pattern fits well.
Wrapping Up
The architecture exists. The hardware is being manufactured. The patterns confidential inference, RAG, federated learning, rack-scale clusters of GPUs, and Kubernetes-native containers are documented and deployable.
Still to be done: standardizing multi-cloud attestation procedures, agentic AI regulation patterns, and more tightly integrating TEEs with the distinct concepts of differential privacy or homomorphic encryption. These are the truly open problems, and the next generation of reference designs will emerge from them.
As a first-time architect, my advice is to stick with one pattern (for attestation and key management, the right first move is to choose one pattern and stick to it). Implement the pattern in a pilot with your preferred cloud offering a confidential VM (the right first move) and get used to attestation and key management before scaling. Learning in operations by the same pilot is worth more than reading ten more whitepapers.
Secret AI is where compliance pressure, hardware capability, and real security meet. Now is that convergence.
I’m a technology writer passionate about AI and digital marketing. I create engaging and useful content that bridges the gap between complex technology concepts and digital technologies. My writing makes the process easy and engaging. I encourage participation I continue to research innovation and technology. Let’s connect and talk technology!



