GDPR for AI customer support: DPA, transfers, retention and data-subject requests
A practical guide to the four questions procurement always asks about an AI support tool: what the DPA has to cover, where the data goes, how long transcripts live, and how you answer an access or erasure request across every channel.
Adding an AI assistant to your support stack does not change a single GDPR obligation. It changes where personal data turns up. A chat box collects what a contact form never does: order numbers, addresses, payment complaints, whole life stories typed at midnight because the customer is annoyed. That content is then retrieved, summarised, sent to a model, stored as a transcript and counted in your analytics. None of it is unlawful. All of it is yours to account for.
1. Who is who, and what each party sees
Most arguments about AI and GDPR get shorter once the roles are said out loud. You decide why and how your customers' data is processed, so you are the controller. Your support platform processes it on your instructions, so it is a processor. The model provider behind the AI answers sits below that as a sub-processor. Three parties work on one chat message, and only one has a relationship with the customer: you.
The chain behind one chat message
Party
Role
What it typically sees
You (the merchant)
Controller
Everything: the conversation, the customer record, the order behind it, the decision about what happens next.
The support platform
Processor
Conversations, contact details, connected order metadata and the knowledge-base content used to answer.
The model provider
Sub-processor
The text sent for one answer — question, retrieved passages, context. Not your database.
Your connected shop or channel
Usually your own controller relationship
Whatever the integration reads, governed by your agreement with that provider rather than your vendor's DPA.
That last row catches people out. A shop platform or messaging channel you connect yourself is a service you directed, not a sub-processor your vendor engaged — and nobody else will put it in your record of processing activities.
2. The DPA checklist
Article 28 of the GDPR — see the text of Regulation (EU) 2016/679 — sets out what a processor contract must contain. Most published DPAs cover it. The useful reading is not whether the clauses exist but whether anything specific sits underneath them.
Subject matter, duration, nature and purpose, and the categories of data and data subjects. Vague here means the vendor has not worked out what it holds.
Processing only on your documented instructions, transfers included, plus a duty to flag an instruction that looks unlawful.
Confidentiality for everyone with access, and security measures described concretely: encryption in transit and at rest, tenant separation, access control, audit logging, backup and restore.
A named sub-processor list, and a notice period before one is added or replaced, with a right to object.
Assistance with data-subject requests and with Articles 32 to 36 — breach handling, impact assessments, prior consultation.
Breach notification without undue delay, with a stated target and a named contact rather than the phrase from the regulation.
Audit rights, and what the vendor offers in place of an on-site audit.
Deletion or return at the end of the contract, including what happens to backups and how long that takes.
Two things that are not clauses matter as much. How you are told about a sub-processor change — a mailing list you must remember to join is weaker than a contractual notice period. And whether your support content trains general models, which should be a plain sentence rather than an inference. Ours says no; the sub-processor tables, the transfer mechanism and the 30-day change notice are in the DPA and the privacy policy.
3. Transfers: where the data sits, and what leaves
"Is it hosted in Europe?" is two questions. Where is the primary store, and which individual operations leave the EEA? For us, primary hosting is in the EU, every workspace's data is stored separately from every other customer's, and it is encrypted in transit and at rest. Some AI processing runs through sub-processors outside the EEA under the European Commission's standard contractual clauses and the UK addendum, written down where you can check it.
Read every other vendor the same way. One that answers "EU hosting" and stops has described its database, not the model call. Ask in writing which processing leaves the EEA, under which mechanism, and whether a configuration exists where it does not happen at all.
4. Retention: the most commonly un-owned data in support
Transcripts accumulate because nobody decided they should not. They are genuinely useful — quality review, content gaps, the awkward moment a customer says they were promised something — and every month of them is also a month of personal data you would have to search, export and erase on request. Storage limitation is not a suggestion.
Set a period per data type. Conversations, AI decision logs and connected commerce data have different useful lifetimes.
Make expiry automatic. A policy that depends on someone running a cleanup is not a policy.
Know what expiry costs. Deleting last year's transcripts also removes your ability to re-check an old resolution, so decide what you keep in aggregate first.
Redact rather than retain. Masking emails and phone numbers on close keeps a thread useful for review and removes most of what makes it sensitive.
State the periods in your privacy notice in the same words you configured them.
In our product these are workspace settings an owner or admin controls: retention for conversations, a separate period for detailed AI decision logs, and a toggle that redacts customer contact details when a conversation is resolved. Whatever tool you use, answer on day one who owns those numbers — the default in most stacks is "keep everything, forever, by accident".
5. Data minimisation, where it actually happens
Minimisation in a chat product is four decisions about what the assistant is allowed to need.
Do not ask for what you will not use. Every pre-chat field is data you now hold. A name and an email answer most tickets; a date of birth almost never does.
Verify by matching, not by disclosing. Tie an order lookup to the email address the customer is chatting from and return a neutral "not found" on a mismatch, so a guessed order number never confirms that someone else's order exists.
Keep the assistant read-only over customer records. An AI that reads an order is a different risk from one that can change it; refunds and order edits belong in a queue for a person to approve.
Know what the widget leaves in the browser. Ours keeps a session identifier in local storage so a returning customer sees their own thread — a fact for your cookie notice, and a fair question to put to any vendor.
Identity matching is the one that becomes an incident when it is missed, and it is where the AI Act and GDPR meet in the same widget. The disclosure side is covered in the EU AI Act chatbot disclosure post.
6. Access and erasure requests, end to end
A subject access request must be answered without undue delay and within one month of receipt, extendable by two further months for complex or high-volume requests if you tell the person inside the first month. Generous — until you try it across four channels and find the same customer three times.
Find every identity. One person may exist as a widget visitor, an email address, a WhatsApp number and a commerce customer id. Decide in advance which identifier you join on; email is usually the only shared one.
Include derived data. Tags, sentiment scores, AI decision logs, analytics rows and mirrored order data are all personal data about that person.
Verify the requester first. An access request answered to the wrong person is a breach dressed as compliance.
Decide what "erase" means in your stack. Deleting rows often breaks financial and analytics integrity, so many systems irreversibly anonymise instead: scrub every identifying field, keep the empty shell. That is defensible under Article 17 as long as the result cannot be re-linked and you can explain it.
Follow the data into indexes and backups. Vector indexes, caches, an export in someone's downloads folder and encrypted backups all outlive the row. Backups are normally left to the rotation to overwrite; say so, and state the window.
Log the request, the verification, the scope and the completion date. A year later that is your only evidence.
Two vendor questions follow. Can you export and erase one person yourself, or does every request become a ticket with the vendor? And is customer content embedded anywhere erasure does not reach? Our route for individuals is on the data-deletion page, and workspace admins can export and erase a single customer from the compliance settings without asking us.
7. Does a support chatbot need a DPIA?
Usually not on its own. An impact assessment is required where processing is likely to result in a high risk to people's rights and freedoms — large-scale systematic monitoring, special-category data at scale, automated decisions with legal or similarly significant effects. A chatbot that answers documented questions and hands off to a person does not normally clear that bar, and a short documented screening is the right output. That changes as you add things: scoring that routes people differently, health or financial detail arriving routinely, an assistant that decides entitlements rather than describing them, call recording. Check your own supervisory authority's mandatory list before concluding you are outside it — those lists are national and not identical.
The email you can send to a vendor
Everything above compresses into eight questions. Send them before a trial rather than after procurement stalls — the answers usually come fast, and slowness is data too.
Where is the primary data store, and which processing leaves the EEA under which transfer mechanism?
Which sub-processors handle our customers' personal data today, and how much notice do we get before that list changes?
Is our support content ever used to train general models, and is that contractual or a setting?
What retention periods can we configure per data type, and does expiry happen automatically?
How does the assistant verify a customer before disclosing order or account data?
Can we export and erase everything about one person ourselves, across every channel, and what does erasure leave behind?
What is your breach notification target, and who is the named contact?
What certifications and audits do you hold today, and what is roadmap?
Do I need a data processing agreement for an AI chatbot?
Yes. Article 28 of the GDPR requires a written processor contract wherever a vendor processes personal data on your behalf, and a chat product does that by definition: visitors type their names, email addresses and order details into it. The agreement should also name the sub-processors behind any AI features, because the model provider sits in that chain too.
Is it a problem if AI processing happens outside the EU?
Not by itself. Transfers outside the EEA are permitted where an appropriate mechanism is in place, most commonly the European Commission's standard contractual clauses plus any supplementary measures needed. What matters is that you can say which processing leaves the EEA and under which mechanism, and that it is written into the data-processing agreement.
How long should we keep chat transcripts under GDPR?
There is no fixed number in the regulation. Storage limitation means keeping personal data only as long as it is needed for the purpose you collected it for, so set an explicit period per data type, make expiry automatic rather than manual, state the same periods in your privacy notice, and decide what you keep in aggregate once the detailed transcripts are gone.
How do I handle a subject access request for support conversations?
Answer without undue delay and within one month of receipt, extendable by up to two further months for complex or high-volume requests if you notify the person inside the first month. The real work is finding every identity that person has across chat, email, messaging and commerce records, including derived data such as tags and AI logs, and verifying the requester before disclosing anything.
Does erasing a customer delete everything, including AI logs and backups?
It depends how the vendor implemented it, which is why it is worth asking. Many systems irreversibly anonymise rather than delete rows, so financial and analytics records survive without identifying anyone, and encrypted backups are normally left to be overwritten in the ordinary rotation within a stated window. Ask specifically whether customer content sits in any search or vector index.
Does a customer support chatbot need a DPIA?
Usually not on its own, and a short documented screening is the right output. A full assessment becomes appropriate when you add automated scoring that routes or ranks people, routinely handle special-category data such as health details, let the AI decide entitlements rather than describe them, or record calls. Check your national supervisory authority's mandatory list, because those lists differ between member states.
Article 50 of the AI Act became applicable on 2 August 2026. What it actually requires from a support chatbot, why the 'obvious' exemption will not save you, and the GDPR questions that were already there.
Live chat installs in about ten minutes — one script tag or a CMS plugin. Here is the full setup: install route, knowledge base, office hours and handover.
Live chat is a person; a chatbot is software. The honest answer is that you do not have to choose — modern widgets run AI first and hand off to a human.
Resolve more tickets automatically.
See how honestly-measured AI resolutions cut your support load — start on the Free plan, no credit card, no sales call to get started.