Little training text
Your language is a fraction of a percent of what large models read. They get the gist and miss the grammar: cases, gender, aspect, the forms that make text sound native.
Large models learn mostly from English. In Croatian, Serbian, Lithuanian or Latvian they write, but not like a native speaker, and they have never seen your terminology. We adapt open models to your language and your domain, and prove the gain on a test set written in your language. Slovenian is our public proof.
Your language is a fraction of a percent of what large models read. They get the gist and miss the grammar: cases, gender, aspect, the forms that make text sound native.
Legal, medical and technical terms in your language are rare online. The model guesses, translates from English, or mixes in a neighbouring language.
Vendors publish English benchmarks. Nobody has measured the model on your documents, in your language, against answers your experts wrote.
How much usable text exists, under which licences, and what has to be built, translated or distilled. An honest answer before anyone trains.
The open model reads your language and your domain until it writes like a native speaker. Only when the language gap is the problem.
Your tasks: answers, extraction, translation, tool calls. Trained on examples from your work, with retention slices so general skills stay.
A test set in your language with answers your experts wrote. Scored before and after, the same way, every version.
The method is the same for any language with enough text. The closest to our public work are the South Slavic languages, which share much of Slovenian's inflection.
Highlighted: closest to our Slovenian work. For every language we start with a data audit and tell you what is realistic.
Any language with enough usable text for the task. Closest to our public work are the South Slavic languages (Croatian, Serbian, Bosnian, Montenegrin, Macedonian), which share much of Slovenian's rich inflection. We also work with Baltic, Central European and other under-served languages, and we assess each one honestly before promising anything.
The evaluation needs native speakers: they write and check the frozen test set and judge free-text answers. They come from your team or are engaged for the project. The training itself is language-independent engineering.
For a narrow task, a few thousand good examples already move the needle. For continued pre-training, hundreds of millions of tokens. The first step is a data audit that tells you what exists, what licences allow, and what has to be built or translated.
Yes. The weights are delivered to you and run on your own servers, in an air-gapped environment, or EU-hosted. MediaAtlas is an EU company under EU law.
Prices follow the public price list. Agents on top of your model: AI agents.
Tell us the language and the task. We'll tell you how much usable text exists, what a model could do, and what it would take.