How AI Companies Are Finally Locking Down Their Models

Home >> TECHNOLOGY >> How AI Companies Are Finally Locking Down Their Models
Share

Last updated on September 19th, 2026 at 08:44 am

It is the silent anxiety that permeates each AI team that has ever deployed a model to the third-party cloud. You have taken months of your life, even years, training something valuable. Then you hand it over to infrastructure you don’t control at all, run by people with administrative rights over it and no connection to you.
It’s not paranoia. It’s a real gap. And confidential computing is the most plausible technical solution the industry has managed to come up with to date.

It then separates what’s working now, what’s still being figured out, and, most importantly, how engineers, enterprises, and model providers can start using this without getting lost in the complexity.

The Problem Nobody Talks About Out Loud

Most AI security discussions focus on encrypting data at rest and in transit. That sounds complete until you learn there is a third state: data in use. When your model is actively executing inference, the weights are stored in memory – accessible by anybody with root access to the host machine.

Cloud providers, hypervisors, and other users of a shared set of GPUs aren’t meant to look. But technically, many of them can.

This is what everyone who is using confidential computing is trying to keep secret. Before reading on, the Confidential Computing Consortium’s Confidential Computing 101 primer provides background on the differences between this and standard encryption.

Protecting Proprietary Models and IP with Confidential Computing – What the Tech Actually Does

The Enclave Idea, Without the Jargon

Trusted Execution Environments (TEEs) comprise carving out a hardware-isolated portion of memory, an enclave, where code is executed encrypted. The host OS or hypervisor can’t even read what’s happening inside. The two primary CPU implementations that you will come across today on large clouds are Intel TDX and AMD SEV-SNP.

With AI workloads, this means you can load your model weights into an encrypted container image. Such an image is, in effect, useless unless a trusted TEE has been established. When deployed, you can perform a check equivalent to hardware and software stack attestation, called remote attestation: whether the hardware and software stack is exactly as you intended before you publish the decryption key.

One thing that has helped me comprehend this pipeline better is looking at the documentation that NVIDIA has put together on this pipeline, and the mental model that has helped me make better sense of this pipeline is this: think of it as a sealed vault that can demonstrate to you that it is a vault even before you hand over the combination.

The flow appears to be approximately the following:

  1. Encrypt weights + code model weights with keys maintained in a Key Management System (KMS) or Hardware Security Module (HSM).
  2. Wrap as a container image (encrypted) – will do nothing in plaintext outside a TEE.
  3. Deploy on confidential VMs or enclaves, with a very small trusted software stack.
  4. Run remote attestation -only when measurements check, the KMS issues decryption keys.
  5. Decrypt within the enclave, infer, and clean up logs of any sensitive result.

The bulk of the operating nuance is in step 4.

What’s Production-Ready Right Now

It is no longer experimental, at least not at the CPU end.

All three major clouds (AWS, Azure, Google Cloud) offer confidential VM SKUs based on Intel TDX or AMD SEV-SNP. You can run containerized AI workloads on these today with comparatively few code adjustments. Even Google Cloud provides much of the attestation plumbing in its Confidential Space product.

Other platforms, such as Fortanix Confidential AI, go a step further: they have built end-to-end pipelines with encrypted model distribution, attested key release, and runtime protection that work across enterprise deployments. I realized they’re particularly geared toward regulated industries such as healthcare and finance, where both the model IP and the patient/transaction data must be safeguarded jointly.

Real rollouts already exist in those areas. Hospitals use federated learning across institutions, and banks use third-party LLMs without exposing customer records to the model provider. These aren’t pilots; they are live in production.

To establish a sound foundation for why the enterprises in this case are taking this direction, the dissection of Top Benefits of Cloud Computing on Business published on the Google Cloud Blog is tangible. It does not engage in the ancient art of hand-waving.

Where Things Get Messy – The Real Challenges

GPU Support Is Early, Not Finished

NVIDIA’s H100 and H200 chips introduced confidential computing support at the GPU level, pushing the TEE boundary into the device’s RAM. In the context of LLM inference, that is enormous – even the most serious models cannot possibly run on CPU TEEs alone due to performance limitations.

However, the software-based ecosystem surrounding GPU secret computing continues to grow. Adapting frameworks such as vLLM or PyTorch is needed to run them correctly in a confidential GPU environment. Debugging within an enclave is deliberately limited by design, which makes it noticeably worse than debugging typical GPU workloads.

A paper published by IBM Research showed that, with the right parallelization strategy, the overhead of CPU-GPU TEE setups can be insignificant relative to LLM use. Yet that sentence does a great deal of work with that phrase, with the right strategy. You should benchmark your particular workload, rather than using general figures.

Attestation at Scale Is an Operational Problem

Remote attestation is clean in a diagram. Practically, you’re tracking the attestation flows, key release by policy, and schedules of key rotation, audit trails, and that must all be integrated with your existing CI/CD pipelines and MLOps tooling.

Based on Google Cloud’s Confidential Space documentation, the primitives will be solid, but wiring up the actual functionality will be complex. This will hit for those who have never considered managing secrets at scale.

Confidential Computing Doesn’t Replace Everything Else

This is the lesson that most articles will omit. Hardware isolation safeguards your model weights and in-memory data from infrastructure-level attackers. It doesn’t protect against:

  • A legitimate user deriving knowledge out of the model by the tactic of repeated queries (model extraction attacks)
  • The first place is weak access controls for inference requesters.
  • Contract variation on deployment to consumers, who do not agree to extraction of the model, but might attempt to anyway.
  • Translation requirements that transcend technical constraints.

Contracts, licensing, access controls, rate limiting, legal frameworks are still there. Confidential computing is one part, but not the entire stack.

What’s Coming And Why It Actually Matters

NVIDIA Blackwell and the Next GPU Generation

NVIDIA’s Blackwell architecture (B100/B200) is expanding confidential computing support even further, with enhanced integration between the CPU TEE and regions of Google compute resources that are provided and handed out protected by confidential computing. The economics of confidential AI inference will change as these chips become more ubiquitous across cloud providers and markets that offer it.

Right now, GPUs in confidential instances are both very costly and scarce. That will change.

Cross-Vendor Attestation Standards

Currently, attestation is mostly vendor-specific. Intel TDX attestation does not inherently or automatically interoperate with the AMD SEV-SNP verification in a standardized manner. The Confidential Computing Consortium has been developing the standards, but they are not yet complete.

Cross-vendor attestation can make it much more straightforward to build portable confidential AI pipelines that aren’t tied to a single cloud or even a single chip vendor. When the ecosystem is truly opened up, this is what happens.

Tooling and Developer Experience

Developer friction is real, and vendors understand it. Improved observability tools that don’t break enclave guarantees, improved debugging workflows, and higher-level SDKs are all areas of active investment. The trend is clearly to make this available to teams that aren’t necessarily confidential computing experts.

Three Scenarios Worth Understanding

These overlap with the actual distribution of this by various teams:

Scenario 1: The Model Provider: A business with a proprietary LLM desires to provide deployments of the bring-your-own data. The encrypted model runs on the customer’s confidential, GPU-based infrastructure. The provider can still protect IP, and the customer can still retain full data sovereignty. No one needs to trust either party’s infrastructure.

Scenario 2 – The Enterprise: A financial services company has standardized on confidential VM and GPU SKUs across all AI workloads that interact with customer data. The platform enforces key release based on attestation. The security team can show auditors that model weights and inference data are not shared with the cloud provider’s administrators.

Scenario 3 – The ML Engineer Getting Started: Begin with a very simple inference workload on a confidential VM with sample code in Google Cloud or Azure. Know the flow of the attestation. Benchmark latency. Once you’re familiar with the patterns, move high-value models into the pipeline.

My Take on Confidential AI Platforms Specifically

Systems such as Fortanix Confidential AI abstract away much of the complexity. For teams that need to distribute encrypted models and protect runtime but don’t want to build an attestation infrastructure by hand, these platforms are worth considering.

What I found interesting is that their approach specifically addresses multi-party deployment, where you have a model provider, a data owner, and a compute provider. Still, they are separate entities with different trust relationships. Coordinating that by hand is very difficult. A platform that provides this end-to-end isn’t just a convenience; it’s risk reduction.

Vendor dependence is the tradeoff. You are entrusting the platform’s software stack, and that stack becomes part of your trust boundary. Good to consider carefully.

Free Resources Worth Actually Using

You can continue digging without spending, and this is a viable reading order:

Start with concepts:

  • What is confidential computing? – the clean overview of What is confidential computing? with real AI examples. Red Hat.
  • The blog by Google Cloud on confidential computing in AI and federated learning – good on patterns of architecture talks in a blog.

Shuffle to AI-specific depth:

  • The article by Red Hat Office of the CTO about how AI inference security improves using confidential computing.
  • NVIDIA developer blog on safeguarding sensitive data and AI models – discusses end-to-end pipelines.

Get into benchmarks:

  • arXiv article on confidential computing on NVIDIA Hopper GPUs – will quantify how much overhead there is in the specific case of LLM inference on NVIDIA Hopper GPUs.
  • IBM Research vLLM in secrecy CPU-GPU enclaves – the benchmarking information here is encouraging.

Hands-on docs:

  • Google Cloud Confidential Space documentation
  • ist Fortanix Confidential AI docs and blog posts

Deal with them in that order: ideas, then architecture, then platform specifics.

Who Should Actually Care About This Right Now

Not everybody has to do so at the moment, but here, perhaps, is a definite setback:

Act now if you’re:

  • A model provider that is deployed to environments of customers or shared clusters of GPUs.
  • A business in health, finance, or government that deals with sensitive information and proprietary models.
  • A team that does inference over rented/third-party-provided infrastructure based on a rented set of GPUs.

Prepare to score in case you:

  • Create AI products that will, at some point, process regulated information.
  • Analysis of cloud providers to long-term AI infrastructure.

Check at the moment whether you are:

  • Starting at the early steps and setting up on your specifically managed infrastructure.
  • Operating on open-source designs in which IP protection is not as much of an issue.

Wrapping Up – An Honest Take

Confidential computing can protect proprietary models and IP in real, practical use on CPUs today. On GPUs, it is nearer – in practice, though still crude in toolmaking and supply.

The technology helps bridge a real gap that encryption could never bridge.

Nonetheless, it is part of a broader security and IP strategy, rather than a replacement for access controls, contracts, and compliance processes.

The teams that receive small workloads today will gain familiarity with the new work processes of attestation and key management, and will take advantage of the opportunity to use confidential GPU availability when it becomes the option of last resort.
Not a far-off horizon. Hardware is already shipped.

Leave a Reply

Your email address will not be published. Required fields are marked *