AI agents with human approval for small companies: autonomy tiers, real examples and price
What an AI agent is and is not, five autonomy tiers, a correspondence agent that waits for approval, what to measure, the AI Act and what a pilot costs.
A small company asks: can an AI agent take over part of our correspondence, orders or back office, without sending something wrong to a customer? And what does that cost?
Short answer: yes, if a person approves what matters. Every kind of action gets an autonomy tier. Reading and sorting can run on their own early. Messages to customers wait for approval until the numbers show the agent gets them right. Quotes, prices and contracts always wait for a person. A pilot for one process is from €6,900 and takes 4–6 weeks. For some jobs a tool at about €20 a month is enough, and we will tell you so.
We are a small team in Sevnica, and we run eight agent systems of our own every day, from a correspondence agent to an agentic CRM. The examples below come from them.
An agent is not the same as an automation
An automation in Zapier or Make follows a rule you wrote: when a form arrives, create a row and send a confirmation. It does the same thing every time. That is its strength.
An agent gets a goal and a set of tools. It reads the input, decides which tools to call and in what order, and ends with a draft or an action. For an e-mail about a damaged delivery it might look up the customer in the CRM, check the order in the ERP, read the relevant clause of your terms and draft a reply in Slovenian.
The difference shows up in the price. If your process fits into "when X, do Y", build an automation. It is cheaper, more predictable and easier to check. An agent pays off where the input is mixed (e-mails, scans, phone calls), where the next step depends on what is inside, and where someone today spends hours reading before they can act.
Five autonomy tiers, from draft to acting alone
In our agentic CRM (9 agents, 99 tools) each action class has an autonomy tier and a risk score. Riskier actions wait in an approval queue with a deadline and escalation. The five tiers:
| Tier | Runs on its own | Waits for a person |
|---|---|---|
| Draft only | nothing | everything, including reading and summaries |
| Low | reading, search, summaries, classification | internal writes and all external messages |
| Medium | plus internal writes: tasks, notes, record updates | external messages |
| High | plus routine external messages: replies, reminders, confirmations | bulk mailing, deal changes |
| Full | plus high-value and bulk actions | quotes, prices, contracts |
Two rules sit above the table. Raising a tier is a human decision, not something the agent does for itself. And quotes, prices and contracts never run alone, at any tier.
This is an example policy. For your company we set the thresholds per action class. Most small companies start at low or medium and stay there longer than they expected. That is fine.
Example: a correspondence agent that waits for approval
Our own correspondence agent sorts the inbox and drafts replies and follow-ups. Nothing is sent until a person approves it in Telegram or the review app, and one command pauses everything. 100 % of its outgoing messages are approved by a person.
Here is an example run at the medium tier, on a claims e-mail, as shown on our agents page:
- read: the e-mail and the attached scan
- tool:
crm.lookupfinds the customer - tool:
policy.checkchecks the clause and the deadline - draft: a reply in Slovenian, 182 words
- risk: external message, score 0.41, needs approval
- human: the claims officer approves with one edit
- send: labelled "AI-assisted, reviewed by a person"
- record: the run is written down with the model version; inputs and outputs are stored as hashes, so personal data is not copied into the log
The record is what makes a tier change possible later. Without it you cannot say how often the officer edited the draft, or why.
What to measure before raising a tier
Three numbers, measured on your cases rather than on a demo:
- Share of drafts approved without an edit. If the officer changes one draft in three, the agent is not ready for the next tier. Look at what they change: tone, facts, or a wrong tool result.
- Tool-call errors. Wrong tool, wrong arguments, or a call that should not have happened. We measure this on a frozen set of held-out steps from your process.
- Correct "no call" decisions. An agent that calls a tool when it should simply answer does damage quietly. In our training flywheel, fine-tuning raised the share of correct "no tool call" decisions from 0.847 to 0.925 on 2,471 held-out steps.
Measure over several weeks, not a few days, and measure again after every change of model, prompt or tools. If a number drops, the tier goes back down.
The AI Act: human oversight and disclosure
This is not legal advice.
Mandatory human oversight under Article 14 of the AI Act applies to high-risk systems, for example in hiring or credit scoring, from 2 December 2027. Most agents in a small company handle correspondence, orders or scheduling and are not high-risk. For them approval is a business choice, not a legal duty. We recommend it anyway, because it is the cheapest way to find out where the agent is wrong.
What does apply is Article 50, since 2 August 2026. People who interact directly with an AI system must be told, unless it is obvious. A chat or voice agent that talks to your customers has to say it is AI. For replies a person has approved, we add a short label such as "AI-assisted, reviewed by a person". It costs nothing and settles the question. The dates are in our note on what companies must do by 2 December 2026.
If we build the agent and you put it into service under your name, you are the provider and we are the supplier. We settle the roles in writing before work starts.
What it costs
From our agents page, in EUR, excluding VAT:
| What | Price |
|---|---|
| Evaluation: fifty of your real cases, written verdict | €1,900, one day, credited against a pilot |
| Agent pilot: one process end to end, tools, tiers, approvals, activity log | from €6,900, fixed, 4–6 weeks |
| Agent retainer: monthly re-score, new cases, tool and rule updates | from €1,200 a month |
If the agent needs a model trained on your data, training follows the training price list.
When a €20 tool is enough
We are not the right choice for every company. You do not need us when:
- the process is a fixed rule. An automation tool is cheaper and easier to check;
- a few people want help writing e-mails. A business subscription to a large model, around €20 per user a month, does that;
- the job takes someone five minutes a week. No agent pays that back;
- you have no examples of the process. Collect fifty first, then there is something to measure.
We are worth the money when volume is high, the input is mixed, data should not leave the building, or you need a record of what the agent did and what a person approved.
Frequently asked questions
What is the difference between an automation and an AI agent? An automation follows a fixed rule you wrote: when X happens, do Y. An agent reads the input, decides which tools to call and in what order, and then drafts or acts. Where the rule is fixed, an automation is cheaper and more predictable.
Can an AI agent act without a person approving? Only for the action classes you raise to a higher tier. Quotes, prices and contracts always wait for a person, whatever the tier.
What should a company measure before giving an agent more autonomy? The share of drafts approved without an edit, tool-call errors on a frozen set of your own cases, and how often the agent correctly decides not to call a tool. Raise a tier only when those numbers hold for several weeks.
Does the AI Act require human approval for agents? Human oversight is a legal requirement for high-risk systems (Article 14). For most agents in a small company, the duty that applies now is Article 50: people who interact directly with an AI system must be told.
What does an AI agent pilot cost? A one-day evaluation is €1,900 and is credited against a pilot. A pilot for one process is from €6,900, fixed, in 4–6 weeks. Keeping the agent measured and current is from €1,200 a month.