← Blog
Documents

AI invoice and contract extraction on your own server: when a SaaS is enough and when you need your own model

When a SaaS for incoming invoices is enough, when contracts and sensitive documents need your own model, how to measure accuracy per field, and how to start.

Tadej Fius · MediaAtlas 8 min read

A finance department, an accounting firm or an in-house legal team asks us the same thing: can a model read our incoming invoices and contracts, put the fields into our system, and do it without the documents leaving our server? And do we need our own model for that, or is a subscription enough?

Short answer: for standard incoming invoices that go straight into accounting, a SaaS service is usually cheaper and faster to start, and you should use one. Your own model makes sense when the documents are not standard (contracts, annexes, delivery notes, scans in poor shape), when they carry personal data or trade secrets that may not leave the building, or when you need to know, per field, how often the extraction is right. Then we measure first and train only where the measurement says it pays.

For standard invoices, a SaaS is usually the right choice

Let us be fair to the other side. Services for incoming invoices, in Slovenia for example Kontiqo or Račun123, are built mainly for invoices, a document type that looks roughly the same everywhere: supplier, tax number, invoice number, dates, amounts, VAT, IBAN, payment reference. They start on the day you sign up, and you pay a monthly fee instead of a project. Before you choose one, check two things: that it connects to the accounting software you use, and where it processes your documents.

There is also a date to keep in mind. Under the Slovenian act on the exchange of electronic invoices (ZIERDED), e-invoices between businesses are mandatory from 1 January 2028 (PwC summary). A structured e-invoice in eSLOG or another EN 16931 syntax does not need to be read by any model; the fields are already in the XML. Reading stays useful for invoices to consumers and from foreign suppliers, which may still come as paper or PDF, the archive of past years, and everything that is not an invoice.

When your own model makes sense

  • Contracts and other non-standard documents. A contract has no fixed fields. Parties, term, notice period, automatic renewal, liability cap, penalties, governing law: each can sit anywhere, in any wording, sometimes in an annex. A service built for invoices does not know what you are looking for. A model trained on your examples does.
  • Documents that may not leave the house. Employment contracts, leases with private persons, invoices that reveal health details, documents with trade secrets. If processing has to stay on your server, a SaaS in someone else's cloud is out, however good it is.
  • Public sector. Many public bodies have internal rules on where documents may be processed and have to show that they follow them. With a model on their own server, "the documents did not leave" can be checked on the network instead of promised in a contract.
  • Your own fields and your own rules. Cost centre from the order text, project code from the description, "does this contract deviate from our template". These are decisions, not OCR. On our decision models page, incoming invoices (pay, hold, escalate) and contract review (risk per clause, 1 to 5) are two such schemas.

Measuring: accuracy per field, not "95 %"

A vendor's "95 % accuracy" says little until you know what was counted. Characters? Fields? Whole documents? Whose documents?

We measure on a frozen set of your documents, with the answers your people would give, kept apart from anything the model learns from. The report has one row per field, and each field has its own rule for what counts as correct:

Field Counted as correct when
Invoice number, tax number, IBAN exact match after removing spaces
Dates same day, whatever the format
Amounts same value to the cent
Supplier, contract parties the right entity, not merely similar text
Contract terms (notice, renewal, cap) the right value, with the clause it came from

Then one number decides how much work stays with people: the share of documents where every field is right. It is lower than per-field accuracy suggests. Plain arithmetic: if each of ten fields is right 98 % of the time and the errors are independent, all ten are right in about 82 % of documents (0.98¹⁰ ≈ 0.82). The report shows each field separately, with the wrong examples quoted, so you can see which field fills the review queue.

We have no public extraction number to show here, and we will not invent one. The sets we build come from a customer's documents and stay with that customer. What we can show publicly is the same measurement method on another task: our fine-tuned Slovene speech-recognition model reaches 21.7 % word error rate on held-out parliamentary speech and 23.4 % across domains, measured on fixed public sets (model card). The report format is in our sample evaluation report.

Architecture: OCR, typed output, a threshold, a person

  1. Text and layout. If a PDF has a text layer, we use it. Scans go through OCR with layout, so tables, stamps and footers stay where they belong.
  2. Extraction into a schema. The model returns JSON against a fixed schema: field names, types, allowed values. There is no free text for the next system to guess at.
  3. Checks in code, not in the model. IBAN check digits, tax number check digit, line items add up to the total, net plus VAT equals gross, dates in a plausible range. A failed check sends the document to review, however sure the model was.
  4. A confidence threshold per field. Above it, the field goes through. Below it, a person confirms or corrects it. We set the threshold from the measured accuracy on your set, together with whoever owns the process.
  5. Human review that feeds back. Corrections become training examples for the next version. Each new version is scored against the previous one on the same frozen set before it replaces it.

Anonymise before the data goes further

Extraction is often the first step, not the last. Contract data goes to a large external model for a summary, invoice lines go to an analytics tool, examples go into training. Before that, personal data can be replaced with placeholders and restored afterwards. Doing this well in Slovenian, where names change with case endings, is a task of its own; how we measure it is in our note on anonymising documents with a local model. Training data gets the same treatment, because what a model learns from can end up in its weights.

How we start

Current prices are on our price list.

Step What you get
One-day evaluation fifty of your documents with gold answers, a frozen set, accuracy per field for up to five models, a written recommendation
Pilot S one document type end to end: schema, training on data you already have, thresholds, deployment, 2–3 weeks
Pilot M when the annotations have to be built first, 4–5 weeks
Monthly maintenance monthly re-measurement on your frozen set; retraining when it pays, billed per cycle
EU hosting, if not on your server your model, your weights

How the two paths compare over time:

SaaS Your own model
At the start depends on the vendor, often little or nothing a fixed-price pilot, once
Every month the subscription, often priced by number of documents maintenance, retraining per cycle when it pays, and hosting if the model is not on your server
Break-even pilot price ÷ (monthly SaaS fee − your monthly cost of the model)

If the SaaS costs less per month than running your own model, the model never pays off on price. Then the reasons to have one are contracts, privacy or control, and you should decide on those, not on the spreadsheet. If you are comparing with a large model's API rather than a SaaS, our break-even calculator works it out for your volume.

When you do not need us

  • Domestic invoices that already arrive as e-invoices. Import the XML; nobody has to read it.
  • Standard invoices, a SaaS that works and nothing sensitive in them. Stay with it.
  • A few dozen contracts a year that a lawyer reads in full anyway.
  • No examples of the correct result. Collect fifty first; without them there is nothing to measure against.

Frequently asked questions

Is a SaaS or our own model better for reading invoices? For standard incoming invoices that go into accounting, usually a SaaS: it is cheaper and starts the same day. Your own model makes sense for contracts, non-standard documents and documents that must stay on your server.

How should extraction accuracy be measured? Per field, on a frozen set of your own documents, plus the share of documents where every field is right. A single percentage without saying what was counted is not a measurement.

Do we still need invoice reading once e-invoices are mandatory in 2028? Less for domestic business invoices, because an e-invoice already carries its fields as XML. It stays useful for foreign suppliers, paper, the archive of past years and contracts.

Can extraction run without documents leaving our server? Yes. We deliver the weights, and the model runs on your server, also without an internet connection.

What does it cost to find out whether it works on our documents? It starts with a one-day evaluation on fifty of your examples, credited against a pilot. Current prices are on our price list.