# Assessing data retention policies at AI vendors | LLMnet

[Skip to content](#lm-inhoud)Network/NL[EN](/en/)[Hubhub.llmnet.nlCompare models by task, language, cost and license.](https://hub.llmnet.nl/en/)[Communitycommunity.llmnet.nlPrompt techniques, patterns and system prompts.](https://community.llmnet.nl/en/)[APIapi.llmnet.nlLLMs robust in software: rate limits, routing, structured output.](https://api.llmnet.nl/en/)[Consultancyconsultancy.llmnet.nlIntroducing AI in an organization, from pilot to production.](https://consultancy.llmnet.nl/en/)[Newsnieuws.llmnet.nlDevelopments in AI, interpreted for the Netherlands.](https://nieuws.llmnet.nl/en/)[Benchmarkbenchmark.llmnet.nlMeasure AI quality yourself, for your own tasks.](https://benchmark.llmnet.nl/en/)[Jobsvacatures.llmnet.nlAI roles, salaries and career paths in the Netherlands.](https://vacatures.llmnet.nl/en/)[Learnleren.llmnet.nlAI concepts in plain language, from beginner to builder.](https://leren.llmnet.nl/en/)[Guidegids.llmnet.nlRun AI privately on your own Mac, PC, NAS or home server.](https://gids.llmnet.nl/en/)[Directorydirectory.llmnet.nlMapping the AI ecosystem: tools, models, companies.](https://directory.llmnet.nl/en/)[Radarradar.llmnet.nlSignals from X, research and communities for indie developers.](https://radar.llmnet.nl/en/)[Appsapps.llmnet.nlReviews of AI apps and open-source repos, with tips for people who build their own.](https://apps.llmnet.nl/en/)[llmnet.nl — main site](https://llmnet.nl/)[](https://x.com/intent/post?url=https%3A%2F%2Fconsultancy.llmnet.nl%2Fdata-retentie-en-ai-leveranciers&text=Retentiebeleid%20bij%20AI-leveranciers%20beoordelen)[](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fconsultancy.llmnet.nl%2Fdata-retentie-en-ai-leveranciers)[](https://www.reddit.com/submit?url=https%3A%2F%2Fconsultancy.llmnet.nl%2Fdata-retentie-en-ai-leveranciers&title=Retentiebeleid%20bij%20AI-leveranciers%20beoordelen)[](#)[](https://x.com/intent/post?url=https%3A%2F%2Fconsultancy.llmnet.nl%2Fdata-retentie-en-ai-leveranciers&text=Retentiebeleid%20bij%20AI-leveranciers%20beoordelen)[](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fconsultancy.llmnet.nl%2Fdata-retentie-en-ai-leveranciers)[](https://www.reddit.com/submit?url=https%3A%2F%2Fconsultancy.llmnet.nl%2Fdata-retentie-en-ai-leveranciers&title=Retentiebeleid%20bij%20AI-leveranciers%20beoordelen)[](#)

# Assessing data retention policies at AI vendors

By Ivo Donker — compiled with AI support (Claude & Gemini) · Last updated: 6 August 2026

When selecting and integrating software solutions, data management is always a crucial part of the risk assessment. When using Large Language Models (LLMs) and other forms of artificial intelligence, however, data retention takes on a fundamentally different weight. The input an organization sends to an external AI service often contains unstructured, highly sensitive information: from internal documentation and intellectual property to personal data and trade secrets.

Every API call or interaction with an AI model sends this data directly to the vendor's infrastructure. Where traditional SaaS applications store data in predictable data files and tables, AI systems process textual and visual input in transient context windows, subsequently processing it through various intermediate layers, logging mechanisms, and storage media. Carefully assessing an AI vendor's retention policy is therefore a necessary step to retain control over company information.

Scope: This guide specifically focuses on auditing and assessing retention periods and processing purposes at external AI vendors. It does not cover broad vendor auditing or carrying out a complete data protection impact assessment. For broader insights, see the guides on [due diligence for AI vendors](https://consultancy.llmnet.nl/en/ai-leveranciers-due-diligence) and carrying out an [AI risk analysis and DPIA](https://consultancy.llmnet.nl/en/ai-risicoanalyse-dpia).

## The fundamental questions for the vendor

In practice, standard product pages, marketing environments, and general documentation from AI providers often fail to provide sufficient clarity about exactly what happens to submitted information. Terms like "secure storage" or "privacy-first" don't cover the specific technical substance. To get an accurate picture, organizations need to ask targeted questions that cover the full lifecycle of a data flow.

An effective assessment requires answers to the following five core questions:

- What exactly is retained? Does storage cover only the raw text of the prompt and the response, or are additional metadata, system instructions, IP addresses, and processing parameters also stored?

- How long does the data remain present? Does the information remain present for the duration of the call's processing time, a fixed period such as 30 days, or until the account is actively canceled?

- Where is the storage location? In which geographic area or data center are the servers located where this data temporarily or permanently resides, and does this also apply to backup copies?

- Who has access to the data? Which authorized employees or automated systems of the vendor have the ability to view the stored input?

- What happens upon termination? Is the data automatically, permanently, and irreversibly deleted from both active databases and backups once the term or contract duration expires?

## Three separate purposes for data retention

One of the biggest pitfalls in evaluating an AI service is the assumption that storage serves only a single purpose. In practice, AI vendors typically use three separate processing purposes for retaining input and output data. Each purpose has its own risk profile and requires a separate assessment.

Processing purpose | 
Description | 
Typical retention period | 
Risk impact | 

Service delivery (inference) | 
Directly processing the call to generate the model's response. | 
Session-bound, up to a few seconds | 
Low, provided the data disappears from the memory buffer after the call. | 

Abuse detection (abuse monitoring) | 
Retaining and checking prompts to detect hate speech, fraud, or harmful content. | 
30 days to several months | 
Medium; data remains present temporarily and is often accessible for review. | 

Model optimization (training) | 
Using customer data to further train future AI models and algorithms. | 
Unlimited or long-term | 
High; intellectual property and sensitive material can leak into the model. | 

### The scope of 'We don't train on your data'

Many AI vendors advertise the firm claim that they don't use customer data to train their models. While this is a crucial condition for business use, it's a misconception to think this automatically makes the data transient. The commitment not to train on data does not, after all, rule out that the data is stored for a long time for abuse detection or system analysis.

When a vendor guarantees it does not train on the submitted input, the logical follow-up question is immediately: "Given that the data isn't used for model training, what specific retention periods and access conditions apply to abuse and security monitoring?" The answer often reveals that the data still remains unchanged on external servers for thirty days.

## Variable retention periods and subscription types

Retention policy is not a fixed given at many AI providers, but is instead tied to the subscription type or account tier purchased. Free consumer accounts and standard self-service accounts typically follow a policy in which data is retained by default and used for training and quality control. Enterprise subscriptions or specific B2B agreements, by contrast, often offer the option to disable training and shorten or fully zero out the retention period for abuse detection.

Because these settings differ per account tier, it is essential to explicitly record the agreed retention period in the business contract. Without this specific contractual anchoring, a change in the general terms and conditions or a change of subscription type could cause the retention policy to change unnoticed. For drafting the right clauses, consult the guidelines on [AI contracts and SLA agreements](https://consultancy.llmnet.nl/en/ai-contracten-en-sla).

## The real-world scenario of human review

An often underexposed aspect of AI data retention is the possibility of human review. AI vendors deploy automated filters to detect potential abuse, harmful input, or violations of the terms of use. When such an automated system flags a specific prompt or interaction, the data flow in question is in many cases forwarded to a review team.

In that scenario, employees or contracted third parties of the AI vendor get direct sight of the submitted text. This is not a theoretical risk, but a structured part of the security policy of many large platforms. When assessing a vendor, one should determine in advance:

- Under exactly which conditions a prompt is escalated to human review.

- What security screening and confidentiality obligations apply to the staff who get access to this data.

- Whether it's possible to contractually exclude human review via an enterprise agreement (a so-called 'no human review' clause).

## Chain processing and subcontractors

AI applications rarely run on isolated infrastructure. Many vendors build their services on top of the APIs of larger AI cloud providers, or use specialized external parties for specific tasks such as vector storage, content moderation, or speech-to-text conversion.

When a chain of multiple parties is involved, the question of data retention multiplies with the number of links. If organization A uses an AI application from vendor B, which in turn uses the infrastructure of provider C and engages party D for moderation, a chain of processing arises. The data passes through the systems of three separate parties.

In such a chain, the retention policy is only as strong as the weakest link. If vendor B promises a retention period of zero days but subcontractor C stores the data for thirty days for log analysis, the overall processing risk is still present. Organizations should therefore request the full list of sub-processors and verify that the retention agreements carry through the entire chain.

## The data processing agreement versus online documentation

In practice, AI vendors often refer to documentation pages, privacy policies, or 'Trust Centers' on their website for their retention policy. While these sources offer valuable technical details, they have a legally dynamic character. The content of a webpage can be unilaterally changed by the vendor at any time without prior notice.

To ensure legal certainty, the agreements on data retention, storage locations, and retention periods must be directly included in the Data Processing Agreement (DPA). Documentation without a contractual basis provides no legal safeguard if the vendor decides to change its policy. The data processing agreement must explicitly state that deviations from the agreed retention policy require written consent. For further privacy aspects, also consult the [GDPR privacy checklist in the knowledge guide](https://gids.llmnet.nl/en/avg-privacy-checklist).

## Logs as a hidden storage location

When assessing data retention, organizations focus primarily on the main database where prompts are stored. A common storage location that gets overlooked here is infrastructural logging. API calls pass through web servers, API gateways, load balancers, and internal monitoring networks before reaching the AI model.

Many of these intermediate layers keep standard log files to support system management, troubleshooting, and network security. These logs often contain the full URI, including headers, and in some cases even the content of the request body. This applies not only to the AI vendor's infrastructure, but also to the organization's own internal network layer.

A thorough retention assessment requires mapping the retention period of these log files as well. For a deeper analysis of setting up internal and external logging, see the article on [audit logging and compliance for APIs](https://api.llmnet.nl/en/audit-logging-en-compliance).

## Reducing risk at the source

Instead of relying on the retention agreements of external vendors, organizations can take measures to limit the amount of sensitive data sent out at the source. The less sensitive data that leaves the organization, the smaller the impact of the vendor's retention policy.

Practical techniques to reduce data risk include:

### 1. Data minimization and cleanup

Analyze in advance what information is necessary for the specific AI task. Remove superfluous details, personal identifiers, and confidential company information from the prompt before it is sent to the API.

### 2. Pseudonymization and anonymization

Replace sensitive fields, such as names, national ID numbers, or specific customer numbers, with automated tokens or generic placeholders. Only after the AI model's output has been received are the original values restored within the organization's own secure environment.

### 3. Separation of data flows

Set up a differentiated architecture in which standard, non-sensitive queries are handled via external public AI APIs, while highly sensitive or confidential categories run through a dedicated, isolated route. For the most critical data, a specific technical configuration can be chosen; see the overview of [zero-data-retention API configurations](https://api.llmnet.nl/en/zero-data-retention-api-configuraties).

## Termination, deletion, and the practice of backups

When a contract with an AI vendor ends, or when a specific retention period expires, the data must actually be deleted. Here, a sharp distinction must be made between making data inaccessible and actually destroying it.

The 'soft delete' principle, where a record in the database receives a 'deleted' status but remains physically present, does not meet the requirements for permanent data deletion. The vendor must be able to demonstrate that the data is permanently erased from active systems.

An additional point of attention is the vendor's rotating backup systems. Data that has been deleted from the active database often remains present for some time in incidental backup copies. The agreement must clearly specify the timeframe (for example, a maximum of 30 to 90 days) within which the data is also automatically overwritten and permanently destroyed in these backup archives.

## The practical questionnaire for vendors

To standardize the assessment process and objectively compare vendors with each other, it is advisable to use a fixed questionnaire. This prevents every evaluation from having to start from scratch and ensures streamlined record-keeping.

An effective questionnaire covers at least the following clearly defined points:

- Is the data sent via the API or interface (prompts, uploads, generated output) used for training, fine-tuning, or improving models?

- What is the exact retention period for data on active servers for, respectively, primary service delivery and any abuse detection?

- Does human review take place on submitted data? If so, under what circumstances, by whom, and can this be contractually excluded?

- In which geographic locations are the data and any derived data stored and processed?

- Which subcontractors or third parties process (parts of) the data flow, and how is their retention policy secured?

- How long does data remain present in backup environments after it has been deleted from the active system?

- Is the specified retention policy guaranteed in a legally binding Data Processing Agreement?

By systematically collecting these answers and including them in the internal record for compliance and risk management, the organization retains control over external data flows and responsible choices can be made when deploying AI technology. For organizations that want to establish an overarching vision, it is advisable to anchor this retention research in the broader process of [drafting AI policy](https://consultancy.llmnet.nl/en/ai-beleid-opstellen).

## Further reading

- [Due diligence for AI vendors](https://consultancy.llmnet.nl/en/ai-leveranciers-due-diligence)

- [carrying out an AI risk analysis and DPIA](https://consultancy.llmnet.nl/en/ai-risicoanalyse-dpia)

- [AI contracts and SLA agreements](https://consultancy.llmnet.nl/en/ai-contracten-en-sla)

- [Drafting AI policy for organizations](https://consultancy.llmnet.nl/en/ai-beleid-opstellen)

- [Zero-data-retention API configurations](https://api.llmnet.nl/en/zero-data-retention-api-configuraties)

- [Audit logging and compliance for APIs](https://api.llmnet.nl/en/audit-logging-en-compliance)

- [GDPR privacy checklist for AI projects](https://gids.llmnet.nl/en/avg-privacy-checklist)

llmnet.nl - B2B AI consultancy and integration
