Decision models · typed answers · EU

Decisions, not prose.

Most of what companies ask a language model is not an essay. It is a choice: which team, which model, allow or block, pay or hold. Decision models answer exactly that. A state and a schema of questions go in, a probability for every allowed option comes out, in one pass. We train them on your cases and run them in the EU.

RoutingGuardrailsTriageEvalsLabelingReal-time
How it works

A schema in, a probability per option out

You describe the decision once: the questions, the allowed answers and what each answer means. The model reads the whole situation and scores every answer to every question together. A new decision is a new schema, not a new project.

State

Texte-mail · chat · ticket
JSON recordtool call · sensor tick
Imagescan · photo · display
Video frames

Decision model

One forward passall questions answered jointly
yes / nochoicescoreimagesyour schema
No generated text, so nothing to parse or validate

Answer

A probability for every option
The top option and its confidence
An expected level for scales
A band: act, confirm or human
Always a valid answerThe model can only pick from the options you defined. No format errors, no retries.
Fast and cheapOne pass instead of a generated reply, so it fits behind every message, row or tick.
MeasurableProbabilities can be checked against your labelled cases and turned into thresholds.
YoursTrained on your cases when needed, run in the EU or on your servers, weights handed over.
Use cases

Nine jobs one model can do

The same model, a different schema. Most systems need two or three of these around the language model they already have.

01

Model routing

Send each prompt to the cheapest model that answers it well. Small questions stay cheap, hard ones get the large model.

prompt→model→small · medium · large
02

Guardrails

Catch prompt injection, abuse and off-policy requests before they reach the LLM, with a score per risk.

input→model→allow · block
03

Tool-call gating

Check what an agent wants to run. Read-only calls go through, money and outside contact wait for a person, destructive calls stop.

tool call→model→allow · ask · deny
04

Inbox triage

Every e-mail, ticket or letter gets an action, a department and an urgency in one pass.

message→model→reply now · later · archive
05

Reranking

Score each passage against the question and keep the ones that answer it, before a model writes anything.

query + passages→model→relevance per passage
06

LLM evals

Grade an answer against its source on your scale and flag claims the source does not support.

context + answer→model→score 1-5 · unsupported?
07

Bulk labeling

Label a whole table in one pass per row, several fields at once.

rows→model→labels per field
08

Real-time control

One decision per sensor tick or game state, fast enough for a control loop.

state→model→action
09

Confidence gate

The probability decides what happens next. Act, ask for confirmation, or hand over to a person.

decision→model→act · confirm · human
Confidence gate

The threshold decides, not the model

Every answer comes with a probability. Three bands turn it into what happens next, and you set the limits from your own labelled cases.

p > 0.9

Act

The decision runs on its own: the ticket is routed, the row is labelled, the call goes through.

0.5 – 0.9

Confirm

The decision is prepared and a person confirms it with one click.

p < 0.5

Human

The model is unsure, so a person decides from scratch and that case becomes training data.

The limits above are a starting point. On your cases we measure how often the model is right in each band, then move the limits until the automatic band is as accurate as the job requires. A model that is too sure of itself shows up here, before it reaches production.

The cases a person decides go back into the evaluation set and the next training round. That loop is the same training flywheel we run for language models.

Where we use it

Decisions in our products and for our customers

Each card is one schema. The numbers in the cards are illustrations, not measurements.

Sensoram

Cold-chain alarm triage

Thousands of readings a day and most alarms are a door left open for a minute. The model reads the reading, the trend and the room, and says what the alarm actually is.

  • real excursion or door
  • stuck sensor
  • limit breach for the report
  • reading from a display photo

For: Pharmacies, hospital stores, labs and food logistics that answer to an inspector.

EDC

Working-time records

People write "sick since Tuesday, child has a fever" and the record needs an absence type. Corrections need someone to approve them.

  • absence type from free text
  • overtime cap at risk
  • correction: approve or ask
  • missing clock-out

For: Employers and payroll bureaus keeping records under the Slovenian working-time act.

AI training

Judge and dataset curation

Training data gets better one decision at a time. The same model that grades answers also decides which rows stay in the training set.

  • keep or drop a row
  • near-duplicate
  • answer quality 1-5
  • unsupported claim

For: Anyone building a fine-tune or an evaluation set from their own records.

Accounts payable

Invoices, reminders, credit notes and the odd phishing mail arrive in one mailbox. Each one needs to go somewhere, and some must not be paid.

  • document type
  • duplicate invoice
  • pay · hold · escalate
  • scan legible

For: Finance teams, accountants and shared-service centres.

Security alert triage

The SOC sees the same alert types every night. The model sorts true from false positives and says whether to isolate a machine now.

  • severity
  • true or false positive
  • isolate · monitor · ignore
  • phishing or not

For: Internal SOCs and managed security providers.

Contract review

Before a lawyer reads a contract end to end, the model flags which clauses deserve the time and how risky each one is.

  • contract type
  • clause risk 1-5
  • which party it favours
  • deadline in the clause

For: In-house legal, procurement and law firms with high volume.

EnergonX

Energy trading compliance

Market messages and trades have to be checked for inside information and suspicious patterns, every day, without a person reading each one.

  • inside information
  • suspicious trade
  • certificate valid
  • report or not

For: Energy traders, utilities and marketplaces under wholesale market rules.

AssetManiac

Market catalysts

A stream of news and on-chain events, most of it noise. The model scores each item as a catalyst and separates rumour from confirmed fact.

  • catalyst strength
  • rumour or confirmed
  • which asset
  • worth an alert

For: Research desks, treasuries and analysts who watch many sources.

Public-sector inbox

Citizens write to a municipality in every possible way. Each letter needs a department, a type and the legal deadline that applies.

  • department
  • complaint or information request
  • legal deadline
  • needs a reply

For: Municipalities, ministries and public agencies.

Mali Pokovci

Rubric grading

A teacher's rubric becomes the schema. The model grades each answer against it and says which idea was missed, so the teacher sees where to help.

  • rubric level
  • concept missed
  • hint needed
  • ready for the next step

For: Schools, course providers and training departments.

Content moderation

Posts, comments and images from users. The model decides what stays, what a moderator should see, and what goes, with the reason as a label.

  • allow · review · remove
  • personal data visible
  • harassment
  • image: minor present

For: Communities, marketplaces and media with user content.

How we build

From your cases to a model you can trust

  1. Write the schema

    Together we turn the decision into questions and options, each with a one-line meaning your team agrees on.

  2. Label real cases

    A frozen set of your own cases with the right answer, kept apart from training. Everything is measured against it.

  3. Measure first

    The untrained model goes through the set, in your documents' own language. If the score is good enough, we stop here and you have the numbers.

  4. Train on your data

    Where the score falls short, for a task or for a language, we train the model on your records, as in our model training service.

  5. Set the gate

    Act, confirm and human limits from the measured accuracy per band, agreed with whoever owns the process.

  6. Run and improve

    On your servers or in the EU. Cases a person decided return to the set; the model is re-scored and retrained when it pays off.

Pricing

Same price list as model training

EUR, excluding VAT. Full detail in the price list.

Is it good enough as is?

Evaluation

€1,900 one day

Your decision as a schema, fifty of your cases labelled, the untrained model measured. A written verdict.

  • NDA before any data moves
  • Accuracy per option and per band
  • Credited against a pilot
In production

Retainer

from €190 / month

The model stays measured: new cases into the set, monthly re-score, retraining when it pays off.

  • Regression alerts
  • Threshold review
  • Billed monthly
FAQ

Frequently asked questions

What is a decision model?

A model that does not write text. You give it a state (text, a JSON record, an image) and a set of typed questions: yes or no, one of several named options, or a level on a scale. It returns a probability for every allowed option, in one pass. There is nothing to parse and nothing to fall out of format.

Why not just ask a large language model to answer in JSON?

You can, and for low volume it works. A decision model is faster and cheaper per decision, it cannot answer with an option that does not exist, and its probabilities can be measured and thresholded. That last part is what lets you decide when a person needs to look.

Do we need training data?

To start, no: a new question is a schema, and the model answers it out of the box. To trust it, yes: we label a frozen set of your real cases, measure the model on it, and train it on your data where the score is not good enough.

Does it work in Slovenian or our language?

We don't assume it does. Before anything goes live we measure accuracy on your own documents, in your language, against a labelled test set. If a language falls short, we fine-tune the model for it, the same way we train language models for under-served languages.

Where does it run?

In the EU, on our own capacity or vetted EU clouds, or on your servers. The weights we train for you are yours.

How do you choose the confidence thresholds?

From your labelled cases, not from the model's say-so. We measure how often the model is right in each probability band and set the act, confirm and human limits so the automatic band meets the accuracy you need.

Bring us one decision.

The one your team makes a hundred times a day. We write it down as a schema, run it on your cases and tell you whether a model can make it.

MediaAtlas d.o.o. · Sevnica, SloveniaNDA by defaultEU jurisdiction