GraphRAG or plain RAG for company documents? When a knowledge graph pays off
When plain RAG is enough for company documents and when a knowledge graph pays off: failing questions, extraction cost, Slovenian inflection, how to start.
A legal department, a construction firm or an insurer holds thousands of contracts, annexes, minutes and e-mails. They are trying retrieval over documents (RAG) or planning it, and then someone mentions GraphRAG. The question we get is always the same: do our documents need a knowledge graph, or is plain RAG enough?
The short answer, in three lines:
- For FAQs, manuals, policies and instructions, plain RAG is enough. The answer sits in one or two passages.
- A graph pays off when the answer lives in the relations: contract ↔ customer ↔ project ↔ complaint, and when a question takes several steps across several documents.
- Decide on a measurement on fifty of your own questions, not on a demo. A graph costs more to build and to maintain, so the difference has to show in the numbers.
What plain RAG is and what graph RAG is
Plain RAG splits documents into passages, searches them by meaning and by keywords, and adds the passages it finds to the prompt. The model answers from what it sees and cites the source.
RAG over a knowledge graph first extracts entities (customers, contracts, projects, people, amounts, deadlines) and the relations between them, and stores them in a graph. At question time the system finds the relevant entities, follows the relations and gives the model a path through the graph, together with the passages each relation came from. Microsoft described the approach in From Local to Global, where the graph also helps with questions about a whole collection, such as which themes keep coming up.
A graph does not replace passage search. In practice we add it on top.
Questions where plain RAG fails
Plain RAG finds passages that resemble the question. The trouble starts when no single passage holds the answer:
- "Which contracts with customers who have an open complaint expire this year?" The deadline is in the contract, the complaint in an e-mail, and the link between them is the customer code.
- "Which projects used the subcontractor who missed the deadline in Krško?" Two steps: project → subcontractor → their other projects.
- "Who on our side signed the annexes to contracts with the same customer?" The signatories are scattered across ten documents.
- "What kinds of complaints were most common last year?" A question about the whole collection, not about a passage.
On questions like these, plain RAG often finds only some of the passages it needs. The model builds an answer from them that sounds convincing and is incomplete. That is worse than "I don't know", because nobody checks it.
Where a graph does not help: "How many days of leave does the policy give?", "How do I reset my password?", "What does Article 7 of the general terms say?". The answer is in one passage. A graph adds cost here and nothing else.
What a graph costs
A graph has three costs plain RAG does not.
Modelling. Someone has to decide which entities and relations count: customer, contract, annex, project, person, deadline, amount. Too small a schema misses relations. Too large a schema produces a graph nobody maintains. This is work with your people, not with code, so we start with one or two kinds of question and a handful of entity types.
Entity extraction. Every document goes through a model that extracts entities and relations into a fixed JSON schema. With a third-party API you pay for each one. With a local model there is no per-token fee and the documents never leave the server, so a smaller model trained on your examples does the extraction. Microsoft itself published a cheaper variant, LazyGraphRAG, because of indexing cost; it moves most of the work to question time.
Maintenance. A new document means a new extraction. An amended contract means corrected relations. A new document type means a wider schema. Every relation has to point to the passage it came from, or you cannot find the wrong ones. Plain RAG just adds the new passages to the index.
Slovenian: inflected names and duplicate entities
This is where English guides go quiet, and for a graph it is the biggest problem.
The same customer appears in documents as "Novak d.o.o.", "Novak d. o. o.", "družba Novak", "družbi Novak", "Novaka" and "naročnik" (the client). If extraction records each form as its own entity, the graph has six customers instead of one, and a question about all contracts with Novak returns a sixth of the answer. The same goes for people ("Ana Kos", "Ani Kos", "ga. Kos"), places ("v Krškem", "Krško") and projects people call by their own nicknames.
What we do:
- The model returns the base form. Extraction writes the name in the nominative, not in the form found in the text. A model learns this from examples; general models often get it wrong in Slovenian.
- Stable identifiers where they exist. Company registration and tax numbers, the customer code from the ERP, the contract number. Where they exist they win over the name.
- Merging with a reason. When we merge two entities we record why. Doubtful cases go to a person instead of being decided silently.
- Measuring duplicates. On a frozen sample we count how many real customers the graph split into several nodes and how many different ones it fused into one.
We describe the same problem in anonymising documents with a local model: a list of names catches the first form and misses the rest.
Our measurements: what we can show and what we cannot
We have no public comparison of plain RAG, graph RAG and a fine-tuned model on the same Slovenian test set, and we will not invent one. The sets we build for customers come from their documents and stay with them.
On your documents the report looks like this, and the measurement fills the cells:
| Setup | Correct answer | Correct source | Multi-step questions | Duplicate entities |
|---|---|---|---|---|
| Plain RAG | – | – | – | no graph |
| Graph RAG | – | – | – | – |
| Tuned model + plain RAG | – | – | – | no graph |
| Tuned model + graph RAG | – | – | – | – |
We split the questions into simple ones (answer in one passage) and connected ones (several documents, several steps). A graph usually adds nothing on the simple ones, so an average over everything hides where it helps.
What we can show publicly is what fine-tuning changes in behaviour: on our public models on Hugging Face, the share of correct decisions not to call a tool rose clearly after training. That is not a graph measurement, but extraction needs the same skill: knowing when not to record a relation. We use a knowledge graph ourselves too: our AI agents in the CRM work with a knowledge-graph memory.
How to decide
| Your case | What to choose |
|---|---|
| FAQs, manuals, policies | plain RAG |
| Content changes weekly, a source is required | plain RAG |
| The answer needs several documents linked (contract ↔ customer ↔ project) | graph RAG |
| Questions about the whole collection (what repeats, where the patterns are) | graph RAG with community summaries |
| The relations already live in the ERP or CRM | a database query, not a graph |
| Strict output format, your terms, tool calling | fine-tuning together with retrieval |
| You do not have fifty checked questions | collect them first |
When to choose retrieval and when to train a model is covered in Fine-tuning or RAG for documents in your language?.
When you don't need us
- You have a few hundred documents and search already returns the right answers. Stay with it.
- Nobody on your side will own the graph schema. A graph without an owner goes stale within months, and wrong relations are worse than none.
- The questions are really database reports ("all contracts expiring in December"). A query in the ERP answers that, not a language model.
How we start
Current prices are on our price list.
| Step | What you get |
|---|---|
| One-day evaluation | fifty of your questions with correct answers, a frozen set, the candidate models measured on your questions, a written recommendation on whether retrieval or a graph pays off |
| Pilot S | a model trained on documents you already have, measured on the frozen set, 2–3 weeks |
| Pilot M | when the schema and extraction examples have to be built first, 4–5 weeks |
| Monthly retainer | new document types, re-extraction, measurement on the same set |
You get the schema, the extraction model, the frozen set and the evaluation scripts.
Frequently asked questions
What is the difference between GraphRAG and plain RAG? Plain RAG finds passages similar to the question and gives them to the model. GraphRAG first builds a graph of entities and relations from the documents and answers along paths that cross several documents.
When does a knowledge graph not pay off? When the answer sits in one passage: FAQs, manuals, policies, instructions. Also when the relations already live in your ERP or CRM, because a database query is cheaper than a graph built from text.
Does GraphRAG work on Slovenian documents? Yes, if extraction returns the base form of names and merges duplicate entities. Without that the graph splits one customer into several nodes, which is why we measure duplicates separately.
Can we build the graph without sending documents to a third-party API? Yes. Entity extraction runs on a local model on your server or on EU hosting, so the documents stay with you.
What does it cost to find out whether a graph pays off? It starts with a one-day evaluation on fifty of your examples, credited against a pilot. Current prices are on our price list.