← Blog
Model training

A small fine-tuned model for customer support in Slovenian: accuracy, frontier fallback, cost

Which tickets suit a model, how many past replies you need, a 4B–12B model with retrieval and a frontier fallback, how we measure it and what it costs.

Tadej Fius · MediaAtlas 8 min read

The head of support at a shop, an internet provider or a repair service asks us: can AI answer our customers in Slovenian, with our terms and without making things up, and what does it cost?

Short answer: yes, for part of the tickets. A small model (4B to 12B) trained on your past replies, with retrieval over your knowledge base and a fallback to a large model when it is unsure, covers the repeated questions. For sensitive tickets it drafts a reply for a person. How large that part is, nobody should guess. It is measured on a frozen set of your tickets.

Which tickets suit a model and which do not

Good candidates repeat and have an answer in some source:

  • "When is the invoice due?", "How do I change my plan?", "Where is my order?"
  • step-by-step instructions from a manual or the FAQ;
  • sorting a ticket into your category and handing it to the right team;
  • a first reply that collects missing details (order number, device model).

Poor or unsuitable candidates:

  • complaints where the customer expects judgement and an apology from a person;
  • compensation, refunds above a limit you set, exceptions to the terms;
  • anything with legal or health consequences;
  • questions your own team answers differently every time. If there is no single right answer, the model cannot learn it and we cannot measure it.

The line is not technical. You draw it, per category, before anything is measured.

Data: how many tickets, cleaning and anonymisation

Tickets and e-mails are closer to training data than PDFs: the customer's message and your team's reply are already a pair. They are not training data yet.

How many. We start with fifty real tickets with replies checked by someone who knows the work. They become a frozen test set nobody trains on. How many pairs training needs depends on how many kinds of tickets you have. A measurement shows it, not a rule of thumb. Quality counts more than quantity: a thousand unchecked replies teach the model your mistakes.

Cleaning. From the ticket history we strip signatures, quoted earlier mail and automatic confirmations. We drop replies after which the customer reopened the ticket, and duplicate pairs. Replies quoting old prices or terms are not memorised: that belongs in the knowledge base, which the model reads with every answer.

A test set from the latest period. Built from random tickets, it would share near-identical examples with the training part, and the score would look better than real use. So we take tickets from the last few weeks that the model never saw in training.

Anonymisation before training. Personal data the model trains on can end up in the weights. Names, addresses, e-mails, phone numbers and contract numbers are replaced before training starts. How we measure that with Slovenian case endings is in our note on anonymising documents with a local model.

The setup: small model, retrieval, fallback and a person

What we build for customer support:

  1. Retrieval over the knowledge base (RAG). Price lists, terms, manuals. What changes stays in documents, not in the model, and the answer cites its source.
  2. A small tuned model, 4B to 12B. Trained on your pairs, it knows your tone, your terms, your categories and your reply format. It runs on your server or in the EU.
  3. Fallback to a large model. When retrieval finds no source, when the ticket fits no known category, or when the answer fails validation (missing source, wrong format), the ticket goes to a large model through an API. Personal data is removed first and restored after the answer. Which data may never leave, you decide on paper. API calls are passed through at cost.
  4. Human approval. A reply reaches the customer only after a person approves it, at least until measurements show a category is reliable. For complaints, refunds and contracts, approval stays. How autonomy is raised step by step is in our note on agents with human approval.

The fallback share is useful in itself. It shows where the small model needs more examples, and it is the first number we track in the monthly retainer.

How we measure it

On your frozen set we compare several setups, all with the same retrieval. The report looks like the table below; the cells are filled by the measurement on your tickets:

Setup Correct answer Correct source Correctly handed to a person Fallback share Cost per 1,000 tickets
Open model, untuned + retrieval – – – – –
Large model via API + retrieval – – – – –
Tuned 4B–12B model + retrieval – – – – –
Tuned model + fallback to large model – – – – –

Free text is scored by an independent judge model against criteria you sign off before the measurement. A person from your team checks a sample. Results are shown per ticket category, because an overall average hides the category that fails. The weakest example is always quoted verbatim.

We have no public customer-support number to show here, and we will not invent one. What we can show is what training changes in the behaviour closest to support: "I know when not to act". On Sraka 27B, across 2,471 held-out steps, correct decisions not to call a tool rose from 0.847 to 0.956. On Slovenian-LLM-Eval it scores 0.871 against 0.844 for its base. The full report format is in our sample evaluation report.

When a model without training is enough

Honestly: often.

  • An open model without training. If your answers are mostly facts from the knowledge base, tone matters little and there are few categories, an open model with good retrieval can be enough. In the evaluation we measure it as one of the candidates.
  • A large model through an API. At a few hundred tickets a month, with good answers and data that may leave the company, stay with the API. The break-even calculator shows when your own model starts to pay off.
  • No checked replies. Collect fifty first. Without them there is nothing to measure.

Training pays off when ticket volume is high, when the model has to write like your team and use your terms, when data may not leave the building, or when a long prompt with all the instructions costs you on every call.

What it costs

Prices from our public price list, in EUR, excluding VAT:

Step What you get Price
One-day evaluation fifty of your tickets, a frozen set, up to five models measured, results per category, a written recommendation €1,900, credited against a pilot
Pilot S a tuned model with retrieval and fallback, on data you already have, 2–3 weeks €3,200
Pilot M when the pairs have to be cleaned and built first, 4–5 weeks €9,000
Monthly retainer new categories, retraining, measurement on the same set, fallback tracking from €690

Large-model calls are passed through at cost. You get the weights, the frozen set and the evaluation scripts.

AI Act: telling the customer is required

Since 2 August 2026, Article 50 of the AI Act applies: people interacting directly with an AI system must be told, unless it is obvious. A chat on your site has to say so in its first message. For e-mails a person approves before sending, we recommend a short label such as "AI-assisted, reviewed by a person". If you put the system into service under your own name, you can be its provider; we are the supplier. The dates are in our note on what companies must do by 2 December 2026. This is not legal advice.

Frequently asked questions

Can AI answer customers in Slovenian? Yes, for repeated questions with a clear source in your knowledge base. For complaints, compensation and anything that needs judgement, let the model draft and a person decide.

How many past tickets do we need? To start, fifty checked question–answer pairs that become a frozen test set. How many training needs depends on how many kinds of tickets you have, and a measurement shows it, not a rule of thumb.

When does a ticket go to the large model? When retrieval finds no source, when the ticket fits no known category, or when the answer fails validation. Personal data is removed before sending, and every such call is logged and counted.

Do we have to tell customers that AI is answering? In a chat, yes: under Article 50 of the AI Act people must know they are interacting with an AI system unless it is obvious. For e-mails a person approves, we recommend a short label.

What does it cost to find out whether it pays off? A one-day evaluation on your tickets is €1,900 and is credited against a pilot. Pilot S on data you already have is €3,200.