What is GraphRAG?

Try it: the Graph RAG playground builds a simple graph from your own document, without a model, and shows which passages it adds to plain RAG. Graphs built by a language model are coming soon.

GraphRAG is a way of doing retrieval-augmented generation that first turns your documents into a knowledge graph, a map of the things mentioned and how they are connected, and then uses that map to answer questions. It was introduced by Microsoft Research in 2024 and is aimed at the questions that ordinary RAG handles badly.

Where plain RAG falls short

Plain RAG retrieves the few passages whose text looks most like the question. That works when the answer sits in one place. It struggles with two kinds of question:

What a knowledge graph is

A knowledge graph has two parts. Entities are the things: products, people, places, policies. Relationships are the links between them, each with a label. One link written out is called a triple: Frame, covered by warranty for, 5 years.

Here is a small graph built from the bike shop handbook used across this site. The highlighted links are the ones that answer "what can I get a refund for?".

part ofpart of5 years2 yearswithin 30 daysneedsfullup to $40BikeFrameComponentsWarrantyReturnAssemblyRefund
Seven entities and eight relationships. Following the two links into Refund finds both the returns rule and the assembly rule, although they are in different sections of the handbook.

How GraphRAG builds the graph

  1. SplitThe documents are cut into passages, as in plain RAG.
  2. ExtractA language model reads each passage and lists the entities and relationships in it.
  3. MergeThe same entity found in many passages becomes one node, with all its links.
  4. GroupClosely linked entities are gathered into communities, like topics.
  5. SummariseThe model writes a short summary of each community.

All of this happens once, before any question is asked. It is the expensive part, because the model is called for every passage and again for every community.

The example, step by step

From three passages of the handbook, the extraction step would produce triples like these:

EntityRelationshipEntityFound in
Bikecan be returned within30 daysReturns
ReturngivesFull refundReturns
Framecovered by warranty for5 yearsWarranty
Componentscovered by warranty for2 yearsWarranty
Assembly by a local shoprefunded up to40 dollarsAssembly

Now ask "What can I get money back for?". Plain RAG looks for passages that share words with the question; the handbook says "refund", not "money back", so with word matching it finds nothing, and with meaning matching it most likely returns only the returns passage. GraphRAG finds the entity Refund, follows its links to Return and Assembly, and brings back both passages. The model can then answer: a full refund for a bike returned within 30 days, and up to 40 dollars towards assembly at a local shop.

Two ways of answering

Local searchGlobal search
Good forQuestions about particular things and how they connectQuestions about the whole collection: themes, summaries, comparisons
HowFind the entities in the question, follow their links, collect the connected passagesRead the community summaries, answer from each, then combine the partial answers
Cost per questionSimilar to plain RAGHigher: many summaries are read for one question

What it costs

When it is worth it

Words you will meet

On TryRAG: the Graph RAG playground finds entities by simple rules and links the ones mentioned in the same sentence, so you can see the idea working on your own text. Coming soon: building the graph with your own model, with named relationships as described above.

More guides

What is RAG? · How to choose a chunk size for RAG · What are embeddings? · What are tokens, and what do they cost? · Hybrid search: words plus meaning · How to test RAG retrieval · Seven common RAG mistakes