Phase 3: RAG & Knowledge Systems

GraphRAG & knowledge graphs for multi-hop reasoning

Advanced ~16 min read
Think of it this way A friendly analogy. Read this if the technical version feels dense. Show Hide

Imagine you're at a giant library with millions of books. If you want to know "What kind of food do pandas eat?", you can probably find a book about pandas, open it up, and get your answer quickly. It's like asking a librarian for a specific topic, and they point you right to a book that has all the info in one place. This works really well for straightforward questions where the answer is in just a few pages next to each other.

But what if you asked something much trickier: "Which famous explorer from the 1500s traveled to a land known for both active volcanoes and a fruit that grows on a tree, and then wrote about it?" No single book will have all of that. You'd need information from books about explorers, fruits, and volcanoes. Simply looking for books that mention "explorer," "volcano," and "fruit" separately might give you lots of books that don't actually link those ideas in a meaningful way. The information is scattered, like having all the puzzle pieces but no idea how they fit to complete the picture.

This is where a "knowledge graph" comes in, making the library super smart! Instead of just books on shelves, imagine every important idea, person, or place (like "Christopher Columbus," "Pineapple," "Mount Vesuvius") has its own special index card. And these cards aren't just put away alphabetically; they're all connected by colored strings, showing how they relate! "Christopher Columbus" might have a string to "discovered America." "America" might have a string to "is home to pineapples." "Pineapples" might have a string to "is a fruit." "Mount Vesuvius" might have a string to "is an active volcano." Now, when you ask that tough question, the computer doesn't just search for keywords. It starts at "explorer," follows the strings, hopping until it finds a complete path of connections that answers your exact question.

So, instead of just getting a jumble of unrelated books, the computer hands you a curated collection of cards and books, all connected by those strings, showing you exactly how one piece of information leads to another to form a complete answer. This means you can ask incredibly complex questions about how different ideas in the world connect – even things that are far apart and seem unrelated – and the computer can trace those links, like following a detailed treasure map, to give you a really smart, put-together answer. You can uncover hidden connections and understand deeper relationships within information.

The core mental model is simple: a chunk-based vector store encodes "what text is near this query," while a knowledge graph encodes "what facts are connected to this concept." A triple like (Ibuprofen, inhibits, COX-2) and (COX-2, regulates, Prostaglandin synthesis) lets you answer "what does Ibuprofen affect in prostaglandin synthesis?" with a two-hop Cypher query. No single chunk contains that answer. The graph makes the relationship traversable by structure, not by hope that both sentences landed in the same retrieved chunk.

Building the graph has two stages. In the extraction stage, you pass each document chunk to an LLM with a prompt instructing it to return a JSON array of {subject, predicate, object} triples. You can use a small, fast model like GPT-4o-mini for this — you're doing classification-style work, not generation. You deduplicate and normalize entities ("IBM" and "International Business Machines" should merge into one node) using fuzzy matching or an embedding-based canonicalization step, then write the triples into Neo4j using the Bolt driver. In the enrichment stage, you attach each triple back to the source chunk ID so you can retrieve the original text when you need it for grounding. The graph stores structure; the vector store stores prose. You need both.

At query time the pipeline looks like this: (1) extract anchor entities from the user query using the same LLM extraction prompt, (2) look up those entities in the graph, (3) run a bounded multi-hop traversal — typically 2 to 3 hops — using Cypher or Gremlin, (4) collect the resulting subgraph as a list of triples or as structured prose, (5) optionally re-rank or prune the subgraph by relevance, and (6) pass the subgraph plus any high-similarity chunks from your vector store to the generator LLM. The generator then has both structured relational context and verbatim source sentences. A real query like "why is Drug A contraindicated for patients with Condition B" traverses (Drug A, metabolized_by, Enzyme X) -> (Enzyme X, deficient_in, Condition B) -> (Enzyme X deficiency, causes, Drug A accumulation). Vector RAG alone would only retrieve this if the exact causal chain happened to appear in one chunk.

The main competing approach is dense-only RAG with larger context windows. GPT-4o with a 128k context can take in dozens of chunks, and sometimes the multi-hop chain assembles itself from sheer volume. The problem: this is expensive per query, latency is high, and the model still makes "lost in the middle" attention errors on very long contexts. GraphRAG pre-structures the reasoning path so the model gets a compact, high-signal subgraph instead of 50 loosely ordered chunks. The tradeoff is that building and maintaining the graph is hard — entities drift, relationships go stale, and extraction quality is imperfect. If your corpus changes daily, graph maintenance becomes a continuous engineering problem. For relatively static corpora (legal codes, scientific literature, internal documentation that changes monthly) the overhead is worth it. For a live customer support knowledge base where articles change constantly, stick with hybrid vector+BM25 search unless you can afford incremental graph updates.

At scale the operational picture changes significantly. At 10 users, running graph extraction synchronously during ingestion is fine. At 10k users, you need an async ingestion worker queue (Celery or a cloud queue like SQS) that writes triples in batches, and Neo4j needs proper index configuration on entity names and relationship types. At 10M users, the graph itself becomes a shared infrastructure component with its own SLO — you need read replicas, connection pooling, query timeouts, and a caching layer (Redis) for hot traversal paths. Triple extraction costs also scale: extracting triples from one million document chunks at roughly 500 tokens each at $0.15/1M tokens (illustrative) adds up to a real ingestion bill. Cache extraction results aggressively; re-extract only when a document version changes. The query-time traversal itself is generally fast if the graph is indexed — bounded 2-hop Cypher queries on Neo4j return in single-digit milliseconds on graphs with millions of nodes. The LLM call to serialize and generate is your latency bottleneck, not the graph query.

Key Takeaways

  • Model entity-relationship triples explicitly so retrieval can traverse logical chains, not just similarity neighborhoods.
  • Extract triples with an LLM at ingestion time; store them in Neo4j or a property graph database.
  • At query time, identify anchor entities, then traverse K hops before serializing the subgraph as LLM context.
  • GraphRAG adds ingestion cost and graph-maintenance complexity; use it only when multi-hop questions are common.

Pro tips

  • Keep predicate vocabularies small and controlled. If you let the LLM invent predicates freely, you end up with 'caused_by', 'is_caused_by', 'causes', and 'led_to' as separate edge types that mean the same thing. Define a closed schema of 20-40 predicates upfront and prompt the extractor to map to that schema.
  • Always store the source chunk ID on every triple edge, not just on the node. When the generator LLM asks for citations, you can walk back from any edge in the retrieved subgraph to the exact paragraph it came from, making attribution straightforward.
  • Bound your traversal depth aggressively. A depth-4 traversal on a dense graph can return thousands of paths and explode your context window. In practice, 2 hops covers most domain multi-hop questions. Add a LIMIT clause and sort by relationship recency or confidence score if you assigned one during extraction.
  • Run a nightly consistency job that checks for orphaned entities (nodes with no edges), duplicate entities with similar names, and triples whose source chunks have been deleted or updated. Stale graph data is harder to debug than stale vector indexes because the model will confidently traverse outdated paths.

Common pitfalls

  • Mistake: Letting the LLM hallucinate entity names during query-time anchor extraction, so the traversal finds nothing. Fix: Fuzzy-match extracted anchor names against the graph's entity index before traversal; if no close match exists, fall back to vector search.
  • Mistake: Serializing the full retrieved subgraph as a flat triple dump, which overwhelms the LLM with noise. Fix: Prune to the shortest paths between anchor entities, and limit total triples passed to the generator to under 30.
  • Mistake: Treating extraction quality as a one-time concern. LLM extractors produce errors consistently on passive voice, negations, and domain jargon. Fix: Spot-check extraction on 50 representative chunks, measure precision, and add domain-specific few-shot examples to the extraction prompt.
  • Mistake: Building the knowledge graph but not keeping the vector store. Fix: Always run both in parallel. The vector store handles queries with no clear anchor entities; the graph handles multi-hop relational queries. Neither alone is sufficient.

When to use GraphRAG vs other advanced RAG approaches

Option Use when Avoid when
GraphRAG (knowledge graph traversal) Questions require connecting 2+ distinct facts across documents; corpus is relatively stable; domain has well-defined entities and relationships (biomedical, legal, financial). Corpus changes daily; questions are predominantly single-hop lookups; team lacks bandwidth to maintain graph extraction pipelines and a graph database.
Hybrid vector + BM25 search Questions mix keyword specificity with semantic intent; corpus is dynamic; you want low operational overhead with a significant recall improvement over pure vector search. Answers genuinely require traversing causal or relational chains that no single chunk contains.
Query decomposition into sub-queries Multi-part questions can be broken into independent sub-questions each answerable by a single chunk; you want to stay within a pure vector RAG architecture. The sub-questions themselves require multi-hop reasoning; relationships between facts are causal or conditional, not just parallel.
Large-context stuffing (128k+ context window) You need a quick prototype without graph infrastructure; corpus is small enough that all relevant chunks fit; query volume is low enough to afford the token cost. Query volume is high (token cost scales linearly); answers are buried deep in a long context (attention degrades); latency SLAs are tight.

Code Example

python
# neo4j-driver>=5.0, openai>=1.0
import os
from neo4j import GraphDatabase
from openai import OpenAI

client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
driver = GraphDatabase.driver(os.environ["NEO4J_URI"], auth=("neo4j", os.environ["NEO4J_PASSWORD"]))

def extract_triples(text: str) -> list[dict]:
    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{
            "role": "system",
            "content": "Extract entity-relationship triples from the text. Return JSON array of {subject, predicate, object}."
        }, {"role": "user", "content": text}],
        response_format={"type": "json_object"},
    )
    import json
    return json.loads(response.choices[0].message.content).get("triples", [])

def write_triples(triples: list[dict], source_id: str):
    with driver.session() as session:
        for t in triples:
            session.run("""
                MERGE (s:Entity {name: $subject})
                MERGE (o:Entity {name: $object})
                MERGE (s)-[r:RELATES {predicate: $predicate, source: $source_id}]->(o)
            """, subject=t["subject"], predicate=t["predicate"], object=t["object"], source_id=source_id)

# Usage
triples = extract_triples("Ibuprofen inhibits COX-2. COX-2 regulates prostaglandin synthesis.")
write_triples(triples, source_id="doc_001")

How this code works

This code’s job is to extract factual relationships, known as "triples" (subject-predicate-object), from unstructured text using an AI model, and then store these relationships in a Neo4j knowledge graph. This process is fundamental for building a knowledge base from text, essential for advanced GraphRAG applications. The setup imports necessary libraries and initializes the OpenAI client and neo4j driver using environment variables for secure access. The extract_triples function sends user text to gpt-4o-mini with a system prompt to identify these triples. The response_format={"type": "json_object"} option is key here; it tells the LLM to guarantee its response is a valid JSON object, which prevents parsing errors. The code then safely retrieves the list of triples from the LLM’s response, defaulting to an empty list if the "triples" key is missing.

The write_triples function takes the extracted triples and a source_id to track the information's origin. It then opens a Neo4j database session. For each triple, it executes a Cypher query using MERGE. This powerful command intelligently creates nodes (like s:Entity and o:Entity) and relationships (like r:RELATES) only if they don't already exist, or matches them if they do. This ensures the graph remains clean without duplicate entities or relationships. The relationship itself stores the predicate and the original source as properties, providing valuable metadata. The example usage demonstrates extracting triples from a sample sentence and then persisting them in the graph with a specific source_id.

Production-grade example

Adds retries with backoff, connection pooling, token cost logging, timeouts, and graceful degradation on graph unavailability.

python
# neo4j-driver>=5.0, openai>=1.0, tenacity>=8.0
import os
import json
import time
import logging
from neo4j import GraphDatabase
from neo4j.exceptions import ServiceUnavailable, SessionExpired
from openai import OpenAI, RateLimitError, APITimeoutError
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type

logger = logging.getLogger(__name__)
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")

client = OpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=15.0)
driver = GraphDatabase.driver(
    os.environ["NEO4J_URI"],
    auth=("neo4j", os.environ["NEO4J_PASSWORD"]),
    max_connection_pool_size=50,
)

@retry(
    retry=retry_if_exception_type((RateLimitError, APITimeoutError)),
    wait=wait_exponential(multiplier=1, min=2, max=30),
    stop=stop_after_attempt(4),
)
def extract_triples(text: str, source_id: str) -> list[dict]:
    start = time.monotonic()
    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[
            {"role": "system", "content": (
                "Extract entity-relationship triples. "
                "Return JSON: {\"triples\": [{\"subject\": str, \"predicate\": str, \"object\": str}]}"
            )},
            {"role": "user", "content": text},
        ],
        response_format={"type": "json_object"},
        temperature=0,
    )
    elapsed = time.monotonic() - start
    usage = response.usage
    logger.info(
        "extract_triples",
        extra={
            "source_id": source_id,
            "prompt_tokens": usage.prompt_tokens,
            "completion_tokens": usage.completion_tokens,
            "latency_ms": round(elapsed * 1000),
        },
    )
    raw = json.loads(response.choices[0].message.content)
    return raw.get("triples", [])

@retry(
    retry=retry_if_exception_type((ServiceUnavailable, SessionExpired)),
    wait=wait_exponential(multiplier=1, min=1, max=10),
    stop=stop_after_attempt(3),
)
def write_triples_to_graph(triples: list[dict], source_id: str) -> int:
    if not triples:
        return 0
    with driver.session() as session:
        result = session.run(
            """
            UNWIND $triples AS t
            MERGE (s:Entity {name: t.subject})
            MERGE (o:Entity {name: t.object})
            MERGE (s)-[r:RELATES {predicate: t.predicate, source: $source_id}]->(o)
            RETURN count(r) AS written
            """,
            triples=triples,
            source_id=source_id,
        )
        written = result.single()["written"]
    logger.info("write_triples", extra={"source_id": source_id, "triples_written": written})
    return written

def traverse_graph(entity_name: str, max_hops: int = 2) -> list[dict]:
    cypher = """
        MATCH path = (start:Entity {name: $name})-[:RELATES*1..$hops]-(end:Entity)
        RETURN [r IN relationships(path) | {s: startNode(r).name, p: r.predicate, o: endNode(r).name}] AS chain
        LIMIT 40
    """
    try:
        with driver.session() as session:
            records = session.run(cypher, name=entity_name, hops=max_hops)
            chains = [record["chain"] for record in records]
        logger.info("traverse_graph", extra={"entity": entity_name, "chains_found": len(chains)})
        return chains
    except ServiceUnavailable:
        logger.error("Neo4j unavailable during traversal, returning empty", extra={"entity": entity_name})
        return []  # graceful degradation: fall back to vector-only retrieval

def ingest_document(chunk_text: str, source_id: str) -> None:
    try:
        triples = extract_triples(chunk_text, source_id)
        write_triples_to_graph(triples, source_id)
    except Exception as exc:
        logger.error("ingest_document failed", extra={"source_id": source_id, "error": str(exc)})
        raise

How this code works

This code is a crucial piece for building a knowledge graph designed to enhance Retrieval Augmented Generation (RAG). It takes unstructured text, extracts structured relationships, stores them in a Neo4j graph database, and can then traverse those relationships to find interconnected context.

The extract_triples function uses an OpenAI model to convert input text into subject-predicate-object triples. A key detail here is response_format={"type": "json_object"}, which ensures the model consistently returns parseable JSON, making it robust. This function also includes retry logic for common API issues. The write_triples_to_graph function then takes these structured triples and saves them to the Neo4j database. A subtle but critical choice is the Cypher MERGE command; it intelligently creates Entity nodes and RELATES relationships if they don't exist, or matches them if they do, preventing duplicates when ingesting information. The ingest_document function orchestrates these two steps. Once data is in the graph, traverse_graph finds related entities by following relationships up to max_hops away using a Cypher MATCH query, enabling sophisticated multi-hop reasoning.

Practice & master

Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.

Exercise

Build a two-hop GraphRAG query function. Given a starting entity name and a Neo4j driver, retrieve all two-hop relationship chains from that entity, serialize them as a numbered list of 'A -> predicate -> B -> predicate -> C' strings, then pass that list as context to an LLM to answer a user question about the entity.

python
# neo4j-driver>=5.0, openai>=1.0
import os
from neo4j import GraphDatabase
from openai import OpenAI

driver = GraphDatabase.driver(os.environ["NEO4J_URI"], auth=("neo4j", os.environ["NEO4J_PASSWORD"]))
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])

def get_two_hop_chains(entity_name: str) -> list[str]:
    # TODO: Write a Cypher query that matches paths of exactly 2 hops from the entity
    # Return each path as a string: "A -> predicate -> B -> predicate -> C"
    pass

def answer_with_graph_context(entity_name: str, question: str) -> str:
    chains = get_two_hop_chains(entity_name)
    if not chains:
        return "No graph context found."
    # TODO: Serialize chains as a numbered list
    # TODO: Call the OpenAI chat completions API with that context + the question
    pass

if __name__ == "__main__":
    print(answer_with_graph_context("Ibuprofen", "What biological processes does Ibuprofen affect?"))

Quick check

  1. A user asks: 'Which enzyme deficiency makes Drug X dangerous?' The answer requires connecting two triples across different documents. Which retrieval approach handles this best?

  2. You extract triples and notice 'causes', 'caused_by', 'leads_to', and 'results_in' are all used as predicates for the same relationship type. What is the most important fix?

  3. Your GraphRAG system returns accurate answers but query latency is 8 seconds. Profiling shows Neo4j traversal takes 12ms and the LLM call takes 7.8s. What should you optimize first?

Self-check: Describe in your own words why vector RAG fails for multi-hop questions, then sketch the three main steps of a GraphRAG query pipeline. Name one scenario where you'd choose hybrid search over GraphRAG, and explain why.