Under-served EU languages · your data · your servers

Your language. Your model.

Large models learn mostly from English. In Croatian, Serbian, Lithuanian or Latvian they write, but not like a native speaker, and they have never seen your terminology. We adapt open models to your language and your domain, and prove the gain on a test set written in your language. Slovenian is our public proof.

−51 %
held-out perplexity after continued pre-training on 1.78 B Slovenian tokens
+4.1
BLEU, Slovenian → English translation, same model before and after
46.8 → 22.3 %
speech-recognition word error rate after fine-tuning
6/7
Slovenian benchmark tasks improved, +3.1 points on average
All public, with model cards:Hugging Faceresults
The gap

Why general models fall short in your language

Little training text

Your language is a fraction of a percent of what large models read. They get the gist and miss the grammar: cases, gender, aspect, the forms that make text sound native.

No domain vocabulary

Legal, medical and technical terms in your language are rare online. The model guesses, translates from English, or mixes in a neighbouring language.

No proof it works

Vendors publish English benchmarks. Nobody has measured the model on your documents, in your language, against answers your experts wrote.

The recipe

The same four steps we used for Slovenian

Data audit

How much usable text exists, under which licences, and what has to be built, translated or distilled. An honest answer before anyone trains.

first week

Continued pre-training

The open model reads your language and your domain until it writes like a native speaker. Only when the language gap is the problem.

when needed

Fine-tuning

Your tasks: answers, extraction, translation, tool calls. Trained on examples from your work, with retention slices so general skills stay.

every project

Frozen evaluation

A test set in your language with answers your experts wrote. Scored before and after, the same way, every version.

every version
Languages

Where the recipe carries over

The method is the same for any language with enough text. The closest to our public work are the South Slavic languages, which share much of Slovenian's inflection.

CroatianSerbianBosnianMontenegrinMacedonian BulgarianSlovakCzechHungarianLithuanianLatvianEstonianAlbanianand others

Highlighted: closest to our Slovenian work. For every language we start with a data audit and tell you what is realistic.

Who it is for

Built for organisations that work in their own language

  • Public administrations and agencies that answer citizens and process documents in the national language
  • Banks, insurers and telecoms with support, claims and contracts in the local language
  • Hospitals and health systems with records, coding and translation in the local language
  • Universities and language institutes building national models and benchmarks
  • Media and publishers who need correct, native text at volume
  • Companies that cannot send data abroad and need the model on their own servers
FAQ

Frequently asked questions

Which languages can you work with?

Any language with enough usable text for the task. Closest to our public work are the South Slavic languages (Croatian, Serbian, Bosnian, Montenegrin, Macedonian), which share much of Slovenian's rich inflection. We also work with Baltic, Central European and other under-served languages, and we assess each one honestly before promising anything.

Do you need speakers of our language?

The evaluation needs native speakers: they write and check the frozen test set and judge free-text answers. They come from your team or are engaged for the project. The training itself is language-independent engineering.

How much text does it take?

For a narrow task, a few thousand good examples already move the needle. For continued pre-training, hundreds of millions of tokens. The first step is a data audit that tells you what exists, what licences allow, and what has to be built or translated.

Can the model run in our country or on our premises?

Yes. The weights are delivered to you and run on your own servers, in an air-gapped environment, or EU-hosted. MediaAtlas is an EU company under EU law.

Prices follow the public price list. Agents on top of your model: AI agents.

Start with a data audit.

Tell us the language and the task. We'll tell you how much usable text exists, what a model could do, and what it would take.

MediaAtlas d.o.o. · Sevnica, SloveniaEU companyWeights delivered