a close up of a computer screen with a map of the world on it

GraphRAG: Retrieval That Reasons, Not Just Searches

The RAG-from-scratch build on this blog uses pgvector to do what almost every RAG tutorial does: embed chunks of text, store the vectors, and retrieve by similarity. That works well for “find passages that talk about roughly the same thing as this question.” It works much less well for questions that need reasoning across several documents at once – “how does this policy change affect the three teams it references” – because vector similarity has no concept of the relationships between the entities inside your documents. That’s the gap GraphRAG is built to close.

What vector RAG actually loses

A vector store treats each chunk as an isolated point in embedding space. If entity A is mentioned in chunk 1 and entity B is mentioned in chunk 40, and the real answer depends on the relationship between A and B, similarity search has no mechanism to connect them unless a single chunk happens to mention both. Multi-hop questions – ones that require following a chain of relationships across several source documents – are exactly where plain vector RAG falls over, because “most similar to the query” and “connected to the query through two other facts” are completely different kinds of match.

What GraphRAG does differently

GraphRAG extracts entities and the relationships between them from your source documents ahead of time, using an LLM pass over the corpus, and builds an actual knowledge graph – nodes for entities, edges for the relationships connecting them. It then goes a step further than a plain graph database would: it clusters related entities into communities and generates an LLM-written summary of each community, so a query can be answered either by walking specific relationship paths or by pulling in a pre-written summary of an entire relevant cluster, whichever the question needs.

  • Indexing – an LLM pass extracts entities and relationships from each document chunk and writes them into a graph
  • Community detection – a graph algorithm (typically Leiden) clusters densely connected entities into communities
  • Summarisation – an LLM generates a natural-language summary for each community, capturing what that cluster of entities is collectively “about”
  • Query – a question is answered either locally (walking specific entities and their direct relationships) or globally (pulling in relevant community summaries for broad, corpus-wide questions)

Trying it with Microsoft’s GraphRAG library

pip install graphrag

# initialise a project pointing at a folder of source documents
python -m graphrag.index --init --root ./my-project

# run indexing - this makes multiple LLM calls per document
# to extract entities, relationships and community summaries
python -m graphrag.index --root ./my-project

# query once indexing finishes
python -m graphrag.query --root ./my-project 
  --method global 
  --query "What are the main risks across all the documents?"

Neo4j offers a more hands-on alternative if you want direct control over the graph schema rather than accepting GraphRAG’s automatic entity extraction – useful once you already know the entity types in your domain (people, products, tickets) and want to define relationships explicitly rather than have an LLM infer them.

The real cost: indexing is expensive

This is the trade-off that matters before adopting GraphRAG for anything: indexing makes an LLM call per chunk for entity extraction, plus further calls for community summarisation, which is dramatically more expensive up front than generating embeddings for the same corpus. It’s also slower to update – adding a handful of new documents to a vector store is a cheap incremental embed; adding them to a knowledge graph can mean re-running community detection and re-summarising affected clusters. GraphRAG earns its cost on corpora where multi-hop, relationship-heavy questions are common; for straightforward “find the passage that answers this” retrieval, plain vector RAG remains both cheaper and perfectly adequate.

Vector RAG and GraphRAG aren’t rivals

In practice, most production systems that need both use them together rather than choosing one: vector search for fast, cheap “find similar passages” retrieval, and a knowledge graph layered on top specifically for the multi-hop and cross-document reasoning queries vector search can’t answer well. If you’ve already built the pgvector setup from this blog’s earlier RAG post, GraphRAG isn’t a replacement for it – it’s the piece worth adding once you notice your users asking questions vector similarity keeps getting wrong.


Leave a Reply