← Blog
Documents

Anonymising Slovenian documents with a local model: accuracy, GDPR and cost

Why regex fails on Slovenian names and EMŠO, how a local 4B–12B model plus rules is measured for recall, what GDPR counts as anonymous, and what it costs.

Tadej Fius · MediaAtlas 8 min read

A court, a notary's office or a document-archiving company asks us the same thing: can a model anonymise our Slovenian documents well enough that we can publish them, share them or send them on, and can it do that without the documents leaving our server?

Short answer: yes, if you combine rules with a small local model and measure the result on your own documents. Rules catch the structured identifiers. A 4B to 12B model trained on your examples catches names, addresses and context that rules miss. What decides whether you can rely on it is recall per data type, measured on a frozen set. Not a demo, not a vendor claim.

Anonymisation or pseudonymisation: what GDPR counts as which

The two words get used as synonyms. Legally they are not.

Pseudonymisation replaces identifiers with placeholders (PERSON_1, ADDRESS_2) and keeps a table that maps them back. Article 4(5) GDPR defines it, and the data stays personal data for whoever holds the table. It is useful when a document has to go somewhere and come back, for example to an external model, and names have to be restored afterwards.

Anonymisation removes the link for good. There is no table, and nobody can reasonably re-identify the person, not even with other data they have. Under Recital 26, anonymous data is outside GDPR. That bar is high: a rare job title, a small village and a date of an accident can together identify someone even with the name removed.

One nuance from 2025. In EDPS v SRB (C-413/23 P), the Court of Justice held that pseudonymised data may not be personal data for a recipient who has no reasonable means to re-identify it, while it stays personal data for the sender. The sender's duties, such as informing people, remain. The EDPB guidelines on pseudonymisation (01/2025) describe how to do it properly. This is not legal advice; your DPO decides which of the two you need.

Why rules and regex fail on Slovenian

Rules work well where the format is fixed. An EMŠO has 13 digits and a check digit, so a rule finds it and checks it. The same goes for tax numbers, IBANs, e-mail addresses and phone numbers. We always use rules for these, because they are cheap and predictable.

Rules stop working where Slovenian grammar starts:

  • Names change with case. "Denis Kotnik" appears later as "Denisa Kotnika", "Kotniku" or only "Kotnik". A list of names from the header catches the first form and misses the rest.
  • Context decides. In a court decision, the judge's name usually stays and the party's name goes. A lawyer, an expert witness or a neighbour who saw the accident are also people. Rules cannot tell a role from a name.
  • Not every address is a residence. "Gosposka ulica" can be where the offence happened, a company's seat or where the defendant lives. Only the last one identifies him.
  • Not every date is a date of birth. A decision has dates of the accident, the hearing and the ruling. Only one of them, if any, is personal.
  • One person, one placeholder. If "Kotnik" becomes PERSON_1 in one paragraph and PERSON_3 in another, the text stops making sense, and a pseudonymised document can no longer be restored.

A list of every Slovenian first name and surname does not solve this. Many surnames are also ordinary words or place names, so the list removes too much and the document becomes unreadable.

A local small model plus rules: the pipeline

What we build for this kind of task:

  1. Text extraction. OCR and layout for scans, so that headers, tables and footers are not lost.
  2. Rules for structured identifiers. EMŠO with check digit, tax number, IBAN, e-mail, phone, registration plates.
  3. A small model for everything with context. A 4B to 12B model trained on your annotated examples marks persons, addresses, dates of birth and other types you define, in all case forms, and returns them as JSON against a fixed schema.
  4. Code, not the model, does the replacing. The model finds entities; a deterministic step replaces them, groups case forms of the same person under one placeholder and writes the key table if you need pseudonymisation.
  5. Measurement on a frozen set. The same documents, the same gold answers, every version.

We measure three things, per data type:

  • Recall: of all personal data in the document, how much was removed. Every miss is a leak, so this comes first.
  • Precision: of everything removed, how much was actually personal. Low precision means an unreadable document.
  • Consistency: whether the same person always gets the same placeholder, and whether restoration from the table gives back the original.

A single average across types hides the problem. A high overall score can sit on top of one type, say dates of birth, that leaks again and again. The report shows each type separately, with the missed examples quoted verbatim.

What the before-and-after numbers look like

We have no public anonymisation number to show here, and we will not invent one. The sets we build for a customer are made from that customer's documents and stay with them.

What we can show publicly is the same method on other tasks. On a Slovene speech-recognition model, fine-tuning brought the word error rate from 46.8 % to 22.3 %. Sraka 27B scores 0.871 on Slovenian-LLM-Eval against 0.844 for its base. Both are measured on fixed sets before and after training. Our sample evaluation report shows the format.

On your documents, the table has one row per data type (person, address, date of birth, EMŠO, other identifiers) and one column per configuration: rules only, an untuned model plus rules, and a trained model plus rules. Each cell holds recall and precision on your frozen set. The recommendation says in writing whether the trained model is worth it and which type still needs a human check.

Why the documents should stay on your server, and what that does to the DPIA

The anonymisation step sees the documents before anything has been removed. That is the most sensitive moment in the whole process. If a foreign API does it, the raw text, with names, EMŠO numbers and health details, goes to a processor and often to a third country.

With a local model:

  • there is no outside processor and no transfer in your records of processing (Article 30) and impact assessment (Article 35);
  • "the data did not leave" can be checked on the network, not only promised in a contract;
  • you still need a legal basis, a retention rule for the originals and for the key table, and access control for both;
  • logs are personal data too. We usually log only the model version and hashes of input and output.

A local model does not remove the need for a DPIA where the risk is high. It makes the assessment shorter and easier to defend. If a pseudonymised document then has to go to a large external model, that step belongs in the DPIA as well. More on that setup in our note on running a private LLM on your own server.

When you do not need us

  • A few contracts a month, anonymised by a lawyer who reads them anyway. Keep doing it by hand.
  • Only structured identifiers (EMŠO, IBAN, tax numbers) in fixed forms. Rules are enough, and any developer can write them.
  • No examples of how your people anonymise today. Collect fifty first; without them there is nothing to measure against.

What it costs

Prices from our public price list, in EUR, excluding VAT:

Step What you get Price
One-day evaluation fifty of your examples with gold answers, a frozen set, up to five models measured, recall and precision per type, a written recommendation €1,900, credited against a pilot
Pilot S a trained model plus rules on data you already have, 2–3 weeks €3,900
Pilot M when annotations have to be built first, 4–5 weeks €12,000
Monthly retainer new document types, retraining, re-measurement on the frozen set from €1,200

You get the weights, the rules, the frozen set and the evaluation scripts. The hardware is yours to choose; for a 4B model the machine you already have is sometimes enough.

Frequently asked questions

Is pseudonymised data still personal data under GDPR? For you, yes: whoever holds the key table can re-identify people. For a recipient without the key it can fall outside GDPR, depending on whether re-identification is reasonably likely for them.

Can regular expressions anonymise Slovenian documents? Only the structured part: EMŠO, tax numbers, IBANs, e-mails and phone numbers. Names and addresses change with case endings and depend on context, so rules alone miss them or remove too much.

Which metric matters most for anonymisation? Recall per data type, because every miss is a leak. Precision matters second: over-redaction makes the document useless. We report both, plus consistency of placeholders.

Does a local model remove the need for a DPIA? No. The anonymisation step itself processes personal data. A local model removes the outside processor and the transfer from the assessment, which makes it shorter.

What does it cost to find out whether it works on our documents? A one-day evaluation is €1,900 and is credited against a pilot. A pilot on data you already have starts at €3,900.