← Blog
Model training

Can you fine-tune a model on customer data under GDPR? EDPB Opinion 28/2024 in practice

What EDPB Opinion 28/2024 means for fine-tuning on customer data: legal basis, anonymisation before training, memorisation tests, DPIA and a DPO checklist.

Tadej Fius · MediaAtlas 8 min read

A support team, an insurer or a law firm wants a model that knows their cases. The cases are full of names, addresses, contract numbers and sometimes health details. The question their data protection officer asks before anyone signs: are we even allowed to train a model on this?

Short answer: yes, if you have a legal basis, use only the data the task needs, and remove or pseudonymise personal data before training starts. A model trained on raw customer records is not anonymous just because it is "only weights". The European Data Protection Board said so in December 2024, and you have to be able to show what you did about it.

This is not legal advice. Your DPO decides; this note gives them something concrete to decide on.

What EDPB Opinion 28/2024 actually says

Opinion 28/2024 on data protection aspects of AI models (PDF) was adopted at the request of the Irish supervisory authority. Three points matter for fine-tuning.

The anonymity test. AI models trained with personal data cannot in all cases be considered anonymous. A model counts as anonymous only if two things are insignificantly likely: extracting personal data of people in the training set directly from the model, and obtaining such data, intentionally or not, through queries. Both are judged against "all the means reasonably likely to be used", by you or anyone else.

Case by case. Supervisory authorities decide per model, based on your documentation. The opinion lists what they may look at: how sources were selected, whether anonymised or pseudonymised data were considered, data minimisation and filtering before training, training choices that reduce identifiability, and tests against attacks such as membership inference, regurgitation of training data, model inversion and reconstruction. If you cannot document that the model is anonymous, the authority can treat it as an accountability failure under Article 5(2).

Legitimate interest is possible. GDPR has no hierarchy of legal bases. If you rely on Article 6(1)(f), you need the three steps: a lawful, clearly stated and present interest; necessity, including whether less data would do; and a balancing test where the reasonable expectations of the people concerned count. A customer who wrote to your support desk does not automatically expect their message to become training data.

The opinion also covers what happens when a model was developed unlawfully. The short version: if personal data stay in the model, the problem travels with it.

Why fine-tuning alone does not make a model anonymous

Language models memorise. Not everything and not evenly, but a rare string that appears several times in the training data, such as a customer's EMŠO in a recurring signature or an IBAN in every invoice reply, can come back word for word when the model is given the right beginning.

Small fine-tuning sets make this worse, not better. Each record is seen several times, so a name that sits next to a diagnosis or a debt in ten tickets has a fair chance of being reproduced.

A model that runs only inside your company lowers the risk, and the opinion explicitly says an internal model may be assessed differently from a public one. It does not remove the risk. Employees can query it, and the weights can be copied.

How we handle it in a project

1. Look at a sample first. Before any training, we list which personal data appear in a sample of your documents, and where: headers, signatures, free text, attachments.

2. Remove or replace before training. Structured Slovenian identifiers are caught by rules: EMŠO with its check digit, tax numbers, IBANs, phone numbers, e-mail addresses. Names and addresses in all case forms need a model trained for the task. How we measure that, per data type, is in our note on anonymising Slovenian documents with a local model. Where the task needs to know that two messages are from the same person, we pseudonymise consistently and keep the key table away from the training data.

3. Test memorisation on a frozen set. After training we run a fixed set of probes: beginnings of training records to see whether the model completes the identifiers, direct questions about people who appear in the data, and planted test records with made-up identifiers that should never come back. The report lists every hit verbatim. The same set runs on every new version, so a retrain cannot quietly make things worse.

4. Write it down for the DPIA. Sources, what was removed and how, the probe results, where the model runs and who can query it. That is the documentation the opinion expects authorities to ask for.

Some knowledge is better kept out of the weights altogether. Prices, contract terms and anything a person may ask you to delete belong in a knowledge base the model reads at answer time. Deleting a document there takes a second. Deleting a person from weights means retraining. More on that trade-off in fine-tuning or RAG.

On-premise or a provider's fine-tuning API: who is the processor?

Provider fine-tuning API Training with us, model on your server
Who holds the training data the provider, often outside the EU us in the EU under a processing agreement, or your hardware only
Processor under Article 28 the provider and its sub-processors us, during training; no outside party once the model runs on your server
Transfers to third countries (Chapter V) often, so you need an adequacy decision or other safeguards none
Who can query the model the provider's infrastructure, every call whoever you allow on your network
Deletion at the end per provider terms written in the processing agreement; you keep the weights

With a hosted API every inference call also sends the prompt out again, so the transfer question is not a one-off. Our security page lists what we sign before we see a document.

A 10-point checklist for your DPO

  1. What is the purpose of the model, in one sentence, and which legal basis covers training for it?
  2. If legitimate interest: is the three-step assessment written down, including people's reasonable expectations?
  3. Which data categories are in the source documents? Any special categories under Article 9?
  4. Which fields does the task actually need, and what was removed before training?
  5. Anonymised or pseudonymised? If pseudonymised, where is the key table and who can open it?
  6. Was the removal step measured per data type, with recall reported?
  7. Was the trained model tested for memorisation, and are the hits listed?
  8. Who is the processor at each stage, and is there a processing agreement?
  9. Do any data or prompts leave the EU, during training or in use?
  10. How will you handle an erasure request: knowledge base, retraining, or both?

When you do not need any of this

  • The task works on documents without personal data, such as product manuals or internal procedures. Train on them.
  • The answers depend on facts that change or must be deletable. Use retrieval over a knowledge base and skip training.
  • You have a handful of documents a month. A person is cheaper than a DPIA.

How we start

Current prices are on our price list.

Step What you get
One-day evaluation models measured on fifty of your anonymised examples, plus a written list of personal data types to remove before training
Pilot S training on data you already have, with memorisation probes in the frozen set, 2–3 weeks; anonymisation of raw documents falls under Pilot M
Pilot M when annotations or the anonymisation model have to be built first, 4–5 weeks
Monthly retainer retraining and re-scoring the frozen set (probes included)

Frequently asked questions

Is a model fine-tuned on personal data automatically anonymous? No. EDPB Opinion 28/2024 says models trained on personal data cannot in all cases be considered anonymous. Extraction of personal data from the model and obtaining it through queries both have to be insignificantly likely, and you have to be able to show that.

Which legal basis do we need to fine-tune on customer data? GDPR has no hierarchy of legal bases, so it depends on the case. Legitimate interest is possible if you pass the three-step test: a real, clearly stated interest, necessity, and a balancing test that takes people's reasonable expectations into account.

Does anonymising data before training solve the problem? It solves most of it. If no personal data go into training, little can come out. The anonymisation step itself is processing of personal data and belongs in your records and, where the risk is high, in a DPIA.

What happens if a customer asks us to delete their data? Deleting a record from a database is easy. Removing it from model weights usually means retraining. This is the strongest practical reason to remove personal data before training, or to keep it in a searchable knowledge base instead.

What does it cost to find out what personal data our documents contain? It starts with a one-day evaluation on fifty of your examples, credited against a pilot. Current prices are on our price list.