# MediaAtlas — AI model training, RAG and on-premises deployment MediaAtlas d.o.o., Kajuhova 11, 8290 Sevnica, Slovenia (founded 2012). One engineering team that trains language models on a customer's own data, builds the retrieval and pipeline around them, and deploys the result inside the customer's infrastructure. The customer receives the weights. Canonical pages: https://mediaatlas.si/ai-training.html (EN) · https://mediaatlas.si/sl/ai-training.html (SL) Price list: https://mediaatlas.si/ai-training-pricing.html (EN) · https://mediaatlas.si/sl/ai-training-pricing.html (SL) Published models and datasets: https://huggingface.co/texdata · https://huggingface.co/MediaAtlas Contact: info@mediaatlas.si · +386 40 42 33 99 · https://cal.com/tfius-tfius ## What we do **Fine-tuning and continued pre-training.** Supervised fine-tuning (chat, extraction, classification, translation, tool calling, reasoning distillation) and continued pre-training on domain or language corpora. Full-parameter (DeepSpeed ZeRO) and QLoRA up to rank 256 on 35B MoE. Identity tuning, sequential re-tunes with retention slices so earlier capabilities are not lost. **Dataset engineering.** Corpus building from a customer's raw material: cleaning, deduplication, quality scoring, licence tracking, deterministic splits, teacher distillation, domain machine translation with frozen terminology, synthetic data where material is missing, exports for SFT, DPO and GRPO. **Evaluation.** Frozen benchmark sets with gold answers, objective metrics (accuracy, BLEU/chrF, WER/CER) plus multi-judge tracks for free text, base-versus-tuned scoreboards, regression flags between versions. Evaluation is also sold on its own: one day, on the customer's documents, with a written recommendation on whether training is worth it at all. **RAG and knowledge-graph RAG.** Retrieval pipelines over the customer's documents: chunking and layout-aware parsing, embeddings, hybrid lexical plus vector search, reranking, citation-grounded answers, freshness and access control per document. Knowledge-graph RAG where relations matter more than passages: entity and relation extraction into a graph, graph-constrained retrieval, multi-hop questions answered over the graph, with provenance on every edge. Fine-tuning and retrieval are combined, not opposed: retrieval for facts that change daily, fine-tuning for behaviour, format, terminology and tool use. **Pipelines and agents.** Document ingestion and processing pipelines (OCR, layout, anonymisation and pseudonymisation, extraction into schemas, validation), structured output enforced by JSON-schema constrained decoding, native tool calling, small-model-first cascades with a frontier fallback, evaluation wired into the pipeline so regressions surface before users see them. **Speech.** Slovene speech recognition (streaming, runs on CPU) and speech synthesis, fine-tuned and deployed the same way as text models. **On-premises deployment.** GGUF (llama.cpp, LM Studio) for laptops and single workstations, vLLM for servers, air-gapped installations where the machine has no network. Deployment ladder from a laptop model to a multi-GPU cluster on the customer's premises, or EU-hosted if the customer prefers. Nothing is required to leave the customer's network; the weights, adapters, datasets and evaluation sets are handed over. ## AI agents Agents for regulated industries that run on a model trained on the customer's data, inside the customer's infrastructure: tools into the customer's systems, autonomy tiers per action class (draft-only to autonomous), approval by a named person, a hash-chained activity log with model version and input/output hashes, a stop switch. Quotes, prices and contracts always need a person. MediaAtlas runs eight such systems itself: a local Slovene voice agent (123 tools), a website-builder agent, an accurate-and-findable content agent (daily audit without a model, fact register checked against sources; audit score 22-92 to 100/100 on 4 sites), a human-approved correspondence agent, an agentic CRM with five autonomy tiers, AssetManiac market-analysis agents over 11 blockchain networks, the training and evaluation flywheel, and disco (open source, https://github.com/tfius/disco), a discovery agent whose threads become SFT/GRPO training data. Pricing: evaluation EUR 1,900, agent pilot from EUR 6,900, agent retainer from EUR 690/month. https://mediaatlas.si/agents.html ## Decision models Small models that do not write text. Input: a state (text, a JSON record, images or video frames) and a schema of typed questions: yes/no, one of several named options with a one-line meaning each, or a level on an ordered scale. Output: a probability for every allowed option of every question, computed jointly in one forward pass, so there is nothing to parse and no answer outside the schema. Uses: routing prompts between small and large models, guardrails before an LLM, gating an agent's tool calls (allow / ask / deny), inbox and ticket triage, reranking passages, grading answers against a source (LLM evals), bulk labeling of tables, real-time control loops, and confidence gates: act above a threshold, ask a person to confirm in the middle band, hand over below it, with the limits set from the model's measured accuracy per band on the customer's labelled cases. Applied to cold-chain alarm triage (Sensoram), working-time records (EDC), training-data curation and judging, accounts payable, security alert triage, contract review, energy trading compliance (EnergonX), market catalysts (AssetManiac), public-sector inboxes, rubric grading (Mali Pokovci) and content moderation. Before anything goes live, accuracy is measured on the customer's own documents and language against a frozen labelled set; where a task or a language falls short, the model is fine-tuned on the customer's data. Runs in the EU or on the customer's servers. Pricing: evaluation EUR 1,900, pilot from EUR 3,200, monthly from EUR 190. https://mediaatlas.si/decision-models.html ## EU AI Act review for SMEs Fixed-price review for companies that already use AI tools, three packages by size: Start EUR 490 (up to 10 people), Standard EUR 990 (10-50), Extended EUR 2,400 (50-250), excl. VAT. Contents: inventory of tools with role (provider or deployer) and risk class, an AI-literacy policy and training log (Art. 4), disclosure texts for chatbots, e-mails and generated media (Art. 50), and a gap list. Technical documentation, not legal advice. Free indicative 2-minute self-check on the page. https://mediaatlas.si/ai-act.html ## Under-served languages The same recipe used for Slovene (data audit, continued pre-training where the language gap is the problem, fine-tuning, frozen evaluation written by native speakers) for Croatian, Serbian, Bosnian, Montenegrin, Macedonian, Baltic and other under-served languages. https://mediaatlas.si/languages.html ## How we work Frozen evaluation set first, then a baseline measurement of untuned models, then training, then the same measurement again — every claim is a number on the same set. Each version is compared with its predecessor and with the open base it started from. Data licences are tracked per source and the customer is told when a combination is not commercially usable. ## Public evidence (all verifiable) - 20+ open models on Hugging Face (4B to 35B MoE) with model cards and evaluation results. - Continued pre-training on 1.78 billion Slovene tokens: perplexity −51 % on a 4B model. - Slovene streaming speech recognition: word error rate 46.8 % → 22.3 %. - Slovene machine translation (EN↔SL) and a medical EN→SL translator; biomedical research model (not for clinical use); models with native tool calling. - Open datasets: Slovene medical SFT and eval sets, translation SFT, pre-training corpus; morphology drills built from Sloleks 3.0 (CC BY-SA 4.0). - On Slovenian-LLM-Eval under an identical harness, our 27B domain model is on par with the public GaMS3-12B-Instruct; the claim is parity, not superiority. ## Typical tasks customers bring Anonymisation and pseudonymisation of legal and administrative documents (with round-trip restoration when a document must be sent to an external model), classification and routing of incoming mail, extraction of fields from claims, contracts, specifications and reports, ICD-10 coding from discharge letters, question answering over internal knowledge with citations, support ticket triage, summarisation of official texts, dictation and transcription in Slovene. ## Commercial terms Public price list. One-day evaluation on the customer's documents €1,900, credited against a pilot. One-week trial under NDA. Fixed-price pilots from €3,200 (data exists) and €9,000 (data has to be built). Full programs from €15,000, quoted after a scoping week. Monthly retainer from €690 for continuous retraining (Watch at €190/month monitors without training; a retraining cycle on demand is €1,490) and re-evaluation. GPU time billed separately when continued pre-training is involved. Prices in EUR, excluding VAT. ## Kratko v slovenščini MediaAtlas uči jezikovne modele na gradivu naročnika (fino uglaševanje, nadaljevano predučenje, izdelava podatkov, evalvacija) in jih postavi pri naročniku: na prenosniku, na eni grafični kartici ali v izoliranem okolju brez omrežja. Poleg učenja gradimo iskalne cevovode (RAG) in RAG nad grafom znanja, cevovode za obdelavo dokumentov (anonimizacija in psevdonimizacija, izvleček v sheme, preverjanje) ter agente s klicanjem orodij. Uteži, nabori in evalvacijski seti pripadajo naročniku; podatki ne zapustijo hiše. Slovenščina je naš javni referenčni primer: več kot dvajset odprtih modelov z objavljenimi rezultati.