Build a Knowledge Graph for RAG: A Small, Testable Example
Model evidence-backed relationships, query a scoped path, and validate provenance before adding extraction models or a graph database.

A knowledge graph for RAG represents entities and relationships so retrieval can answer questions about connected facts. Start with a small set of verified relationships, preserve the source for each edge, and test the query path before adding automated extraction.
You do not need a graph database to validate the first relationship model. This tutorial uses SQLite to make a two-hop example runnable locally. Larger or more complex traversals may justify a graph database, but the evidence and identity requirements remain.
Define the question and the evidence
Suppose the question is: which supplier provides a component used by product P1? Represent two relationship types: a supplier SUPPLIES a component, and that component is USED_IN a product. Each edge has a tenant, a source ID, and an active flag.
The fictional records below demonstrate the two-hop query.
import sqlite3
connection = sqlite3.connect(":memory:")
connection.execute("""CREATE TABLE edges (
tenant TEXT NOT NULL, subject TEXT NOT NULL,
relation TEXT NOT NULL, object TEXT NOT NULL,
source TEXT NOT NULL, active INTEGER NOT NULL,
PRIMARY KEY (tenant, subject, relation, object, source))""")
connection.executemany("INSERT INTO edges VALUES (?, ?, ?, ?, ?, ?)", [
("demo", "supplier:S1", "SUPPLIES", "part:C1", "doc:contract", 1),
("demo", "part:C1", "USED_IN", "product:P1", "doc:bom", 1),
("other", "supplier:S2", "SUPPLIES", "part:C1", "doc:private", 1),
])
def suppliers_for_product(tenant, product):
return connection.execute("""
SELECT DISTINCT a.subject, a.source, b.source
FROM edges a JOIN edges b
ON a.tenant = b.tenant AND a.object = b.subject
WHERE a.tenant = ? AND b.object = ?
AND a.relation = 'SUPPLIES' AND b.relation = 'USED_IN'
AND a.active = 1 AND b.active = 1
ORDER BY a.subject, a.source, b.source
""", (tenant, product)).fetchall()
assert suppliers_for_product("demo", "product:P1") == [
("supplier:S1", "doc:contract", "doc:bom")
]
assert suppliers_for_product("other", "product:P1") == []
The join requires both edges to belong to the same tenant. Without that condition, shared entity names can accidentally create paths across customers. A caller still needs authorization to request the tenant; a SQL predicate alone does not authenticate them.
Keep entity identity stable
Use source-backed identifiers when possible. Two suppliers can share a display name, and one supplier can change its name. Avoid merging entities solely because a model produces similar text.
Record aliases and reconciliation decisions separately from the canonical identifier. When extraction is uncertain, keep the uncertain relationship out of high-consequence answers until it has appropriate support. More edges are not automatically better evidence.
Return the path and its sources
The query returns the supplier and both supporting source IDs. The answer-generation step should receive those sources or their authorized excerpts, so it can explain the relationship rather than merely assert it.
Check permission and source validity before providing the evidence to the model. If a source becomes inaccessible, the application should not continue exposing its derived relationship simply because the edge remains in the graph.
Handle corrections and deletion
The active flag is a minimal current-state mechanism. Marking an edge inactive removes it from this query, but it does not delete the underlying row or satisfy every retention requirement. A complete system needs explicit source-deletion and derivative-cleanup behavior.
For historical questions, add effective times and define the query's reference time. Keep ingestion time separate: learning about an old contract today does not mean the contract became effective today. The temporal-memory guide provides an interval-query example.
Add extraction and a graph backend only when useful
An extraction model can propose entities and edges from documents. Validate relation types, identifiers, evidence spans, and source permissions before accepting those proposals. Treat the extracted graph as a maintained representation of evidence, not an independent authority.
If moving to Neo4j or another graph database, port the same acceptance cases to its real driver and query language. Include the two-hop result, cross-tenant separation, missing paths and inactive evidence in those tests. Evaluate whether this path actually improves answers compared with the vector and graph retrieval alternatives before expanding the graph.
For relationships between remembered facts, explore Supermemory’s memory-graph documentation and check whether its update and extension semantics fit your agent. Keep the custom domain-graph example above as a separate baseline for explicit relationship queries.