Every enterprise AI deal now hangs on a deletion promise no vendor can actually demonstrate; you can’t audit an absence. Payments solved this twenty years ago: stop promising to delete sensitive data, and architect so no one downstream ever holds it.
Somewhere right now, a security team is redlining an AI contract. The clause under negotiation is zero data retention (ZDR): the model provider commits not to store, log, or train on the customer’s inputs. ZDR has become table stakes in enterprise AI. All the major model providers now offer versions of it, because no bank, insurer, or health system will ship production data into a black box without it.
But there’s a fundamental challenge with ZDR. A retention clause is a promise about the absence of something. And absences are nearly impossible to prove.
Proving a negative is miserable work
Somewhere right now, a security team is redlining an AI contract. The clause under negotiation is zero data retention (ZDR): the model provider commits not to store, log, or train on the customer’s inputs. ZDR has become table stakes in enterprise AI. All the major model providers now offer versions of it, because no bank, insurer, or health system will ship production data into a black box without it.
But there’s a fundamental challenge with ZDR. A retention clause is a promise about the absence of something. And absences are nearly impossible to prove.
Ask any infrastructure engineer what “we deleted it” actually entails. Data replicates: hot storage, backups, snapshots, disaster-recovery copies, debug logs, monitoring pipelines, and caches. Derived artifacts multiply: embeddings, evaluation sets, fine-tuning corpora, and the growing memory of long-running agents. A deletion commitment has to chase every one of those copies, forever, across systems owned by different teams, and then convince an auditor it caught them all.
That’s why ZDR is enforced today only the way it can be: contractually. SOC 2 reports, annual audits, and indemnification. These are real controls, and serious providers take them seriously. But they verify the process after the fact. They do not, and cannot, demonstrate that a specific customer’s records no longer exist anywhere in a provider’s estate. Every enterprise security team knows this, which is why AI procurement in regulated industries moves at the speed of legal review rather than engineering.
The strongest deletion guarantee is never having to delete anything.
Invert the problem: prove non-possession
There’s an older, better answer to this problem, and the payments industry has been running it in production for two decades. When PCI DSS made possession of card numbers a liability, the industry’s response wasn’t better deletion. It was tokenization: replace the sensitive value with a surrogate before it enters your systems, so the systems that do the work never possess the thing that carries the risk. Auditors don’t need to ask those systems to prove they deleted card numbers. The numbers were never there.
Apply the same move to AI. If sensitive fields, such as card numbers, account numbers, SSNs, names, and medical record numbers, are detected and tokenized before a payload crosses into an AI provider’s infrastructure, the retention question changes shape entirely:
The provider can log, cache, retain, and hold nothing sensitive. There is no deletion to attest, because there was no possession. ZDR stops being a contractual behavior verified annually and becomes an architectural property, verifiable at the boundary on every request.
The convenient asymmetry
The obvious objection: doesn’t stripping data out cripple the model? For the workloads enterprises actually run, the answer is mostly no, and the reason is an asymmetry worth highlighting. The fields that carry the most liability are usually the ones that carry the least signal.
A model that summarizes support tickets, drafts dispute responses, or reconciles transactions learns from patterns of language and behavior averaged across thousands of interactions. It learns little useful from the fact that a specific card number was 4242 4242 4242 4242. Identifiers are, for most inference and analysis tasks, noise with a breach notification attached. Tokenize them (deterministically, so the same customer maps to the same token and the data keeps its shape and linkage), and you remove most of the risk while keeping nearly all of the utility. Where an authorized system genuinely needs the real value back, detokenization restores it downstream, under policy, with an audit trail.
What we should be honest about
A guarantee is only as strong as the detection in front of it. Structured fields are the clean case: a card number in a JSON payload gets caught, full stop. Unstructured text is much harder; an SSN buried in a rambling support transcript is a challenging recall problem, and solutions that guarantee perfect detection are overpromising. The honest claim is different, and we think it’s actually the stronger one: tokenization converts an unmeasurable risk (what did the model memorize?) into a measurable one (what did detection catch?). Detection recall can be tested, benchmarked, and improved from release to release. Model memorization can’t be audited by anyone, including the people who trained the model. Security teams prefer the risk they can measure. We built our business on that preference.
Where VGS sits in this
VGS has spent nearly a decade as the neutral custody layer for the world’s most regulated data class: payment credentials. We operate PCI DSS Level 1 certified infrastructure, achieved ISO/IEC 27001 certification in 2026, and we’re the only non-PSP with direct connections to all four major card networks for network tokens. Hundreds of companies, such as banks, marketplaces, fintechs, platforms, route sensitive data through VGS precisely so their own systems, and their vendors’ systems, never possess it.
Our AI Data Firewall extends that same architecture to AI traffic: a proxy layer between your applications and AI systems that automatically detects sensitive data in structured and unstructured payloads, applies policy (tokenize, mask, or block) before inference, safely reconstructs values for authorized systems on the way back, and logs every event for audit. It’s the pattern described in this post, running today, on infrastructure that was securing regulated data before “prompt” was a noun.
The pattern we see across customers is consistent: an AI initiative with clear ROI, blocked for months in security review, unblocked in weeks once the review question changes from “audit the model provider” to “audit the boundary.” That’s not a sales line; it’s a structural consequence of moving the guarantee to a place where evidence exists.
An open invitation to AI companies
If you build models or AI platforms, every regulated enterprise deal you’re in right now includes some version of the retention negotiation. Your sales cycle is absorbing the cost of proving a negative. There’s a better division of labor. You should not be in the business of custodying your customers’ raw identifiers; it expands your liability surface and slows your deals. A neutral, certified tokenization layer in front of your API makes your ZDR commitments architecturally true, shrinks what your BAAs and DPAs have to carry, and turns your hardest security-review questions into someone else’s audited infrastructure.
We’ve spent a decade being that solution for payments. The AI stack is next.
Building AI for regulated industries?
Talk to VGS about putting verifiable data protection in front of your models.
Learn more



