Free · runs in your browser · 12 questions

Is your data ready for fine-tuning?

The question we hear most is how many examples it takes. Short answer: fifty real examples with checked answers are enough to measure whether training pays off, and for a narrow task with a strict format also enough for a first trial model. Whether you need training at all depends on the task, which the check below works out from your answers.

Free · indicative

Twelve questions, about three minutes

The result says RAG, SFT, CPT + SFT or not yet, with the reasons, how much data preparation to expect and the next step with its price. Your answers stay in your browser. No cookies.

Answer the questions and click "Show result".

    The result is indicative. It follows the same questions we ask on a first call; the evaluation on your own examples gives the measured answer.

    How we read your answers

    How many examples do you need?

    Less than most people expect to start, more than most people expect to finish. We always begin with fifty real examples of the task with answers someone who does the work has checked. We freeze them as a test set that nobody trains on. On that set we measure untuned models first. If one of them already does the job, you do not need training, and you found out for the price of one day.

    For a narrow task with a strict format, such as pulling six fields from an invoice or sorting tickets into your categories, fifty examples are also enough for a first trial model. That is how our one-week trial works. Broader tasks, free-form answers or a new language need more. How many more is something the measurement on your fifty tells you, not a rule of thumb.

    The ceiling is set by quality, not quantity. For our Slovenian medical research model we first built 140,063 instruction records and passed every one through a quality gate before any training started. A thousand examples with unchecked answers teach the model your mistakes.

    PDFs are source material, not training data

    "Can we just fine-tune on our PDFs?" is the second most common question. Not directly. A model learns from pairs: this input, this answer. A PDF is neither. You have two options:

    • Retrieval (RAG): the PDFs are indexed and the relevant passages go into the prompt. Good when the answer is a fact from the document.
    • Training: the text is extracted (OCR for scans, layout for tables) and turned into pairs, for example a document and the JSON it should yield. Good when the model has to learn a format or a way of working.

    Tickets and e-mails are closer to training data: the customer's message and the reply your team sent are already a pair. They still need cleaning, removal of personal data and a check that the replies were good ones.

    Four possible results

    • RAG. Content that changes weekly, answers that need a source, access that differs per document. Training does not help here. See Fine-tuning or RAG for the full comparison.
    • SFT. Behaviour: a strict output format, your terminology, a smaller language, tool calls. In our public tool-calling fine-tune, correct "no call" decisions went from 0.847 to 0.925 on 2,471 held-out steps. Retrieval cannot give you that.
    • CPT + SFT. The language or domain is new to general models and you have a large archive of text. The model first reads the archive (continued pre-training), then learns the task. After continued pre-training on 1.78 billion Slovenian tokens, perplexity of our 4B model dropped by 51 %. This is program-sized work.
    • Not yet. No examples with correct answers. Neither training nor a fair comparison is possible. Collect fifty first.

    Model cards and results for the numbers above are public on Hugging Face.

    When you do not need us

    • You send a few hundred requests a month to a large API, the answers are good, and the data may leave the company. Keep the API. The calculator on the price list shows when your own model starts to pay off.
    • Nobody in the company can say what a correct answer is. Then no one can check the model either, ours or anyone else's.
    • You need a guarantee for a clinical or legal decision. A tuned model does not give one. It can prepare a draft for a person who decides.

    What the next step costs

    If the check says "not yet", the next step is free: collect fifty examples. Otherwise it is one of two:

    • Evaluation, €1,900, one day. Up to five models on a frozen set from your examples, a results table and a written recommendation on whether to train and what. The fee is credited against a pilot. Sample report.
    • Pilot S, €3,900, 2–3 weeks. When you already have checked examples and a test set: one fine-tuned model, a base-vs-tuned scoreboard, a GGUF build and a go/no-go recommendation.

    If we have to build the data, that is Pilot M (€12,000, 4–5 weeks). Programs with continued pre-training start at €20,000. All prices in EUR, excluding VAT, on the price list.

    FAQ

    Frequently asked questions

    How many examples do I need to fine-tune an LLM on company data?

    To measure whether training is worth it: fifty real examples with checked answers, frozen as a test set. For a narrow task with a strict format, that is also enough for a first trial model. Broader tasks need more, and the evaluation on your fifty examples tells you roughly how many.

    Can I fine-tune a model directly on PDFs?

    Not directly. PDFs are source material: for facts from them, use retrieval (RAG); for training, the text has to be extracted (OCR for scans) and turned into input-output pairs, for example a document and the fields it should yield.

    How much does it cost to fine-tune an LLM on company data?

    At MediaAtlas a one-day evaluation on your documents is €1,900, credited against a pilot. Pilot S on data you already have is €3,900 fixed (2–3 weeks), Pilot M where we build the data is €12,000 (4–5 weeks), full programs start at €20,000. EUR, excluding VAT.

    When is RAG enough and fine-tuning not needed?

    When the content changes often, answers must cite a source or access differs per document, and the model already answers in the right format. Then good retrieval is the simpler option.

    Are my answers stored or sent anywhere?

    No. The check runs entirely in your browser, without cookies. Nothing is stored or sent.

    Fifty examples, one measured answer.

    Send fifty real examples of the task. In one day you know whether training pays off, which model, and what retrieval would add.

    MediaAtlas d.o.o. · Sevnica, SloveniaNDA by defaultData stays in the EU