Conversational Memory in LangChain: History, Stores, and Retrieval
Understand LangChain conversation persistence, control context growth, and connect long-term user context without confusing it with thread history.

Conversational memory in LangChain involves two decisions: where state is persisted, and which parts are sent to the model on the next turn. Keeping a transcript in storage does not mean the entire transcript should be inserted into every request. Conversely, trimming a request does not have to delete its source history.
Current LangChain agents build on LangGraph's execution and persistence model. Use the APIs that match your installed version instead of mixing older buffer-memory examples with a current agent constructor.
Start with thread history
The LangChain short-term memory guide describes checkpointer-backed thread state. This example uses a deterministic fake chat model from langchain-core so the test exercises the framework's history handling without a network request.
from langchain.agents import create_agent
from langchain_core.language_models.fake_chat_models import FakeMessagesListChatModel
from langchain_core.messages import AIMessage
from langgraph.checkpoint.memory import InMemorySaver
model = FakeMessagesListChatModel(responses=[
AIMessage(content="Recorded for this demo."),
AIMessage(content="Second demo reply."),
])
agent = create_agent(model=model, tools=[], checkpointer=InMemorySaver())
config = {"configurable": {"thread_id": "demo-thread"}}
agent.invoke({"messages": [{"role": "user", "content": "Project: Cedar"}]}, config)
result = agent.invoke({"messages": [{"role": "user", "content": "Continue"}]}, config)
assert len(result["messages"]) == 4
assert result["messages"][0].content == "Project: Cedar"
The model returns canned text. The assertions check that messages accumulate in the same thread. InMemorySaver is a local demonstration backend; choose durable storage and test process restarts for a deployed service.
Trim the request without breaking its structure
An agent transcript may contain system messages, user messages, assistant tool requests, and tool results. Removing arbitrary messages can leave a tool result without its corresponding call. Apply the trimming mechanism supported by your LangChain version and preserve valid message ordering.
Select the budget using the actual model tokenizer and input limit. Count system instructions, tool schemas, retrieved documents, conversation messages, and reserved output. A fixed number of messages is a rough control because message sizes vary widely.
For a concrete budget, see context engineering. Allocate room for instructions, history, retrieved evidence and the expected output.
Summaries need an evidence path
Summaries help carry decisions across long threads but can lose dates, negations, and identifiers. A handoff summary should preserve confirmed decisions, unresolved questions, relevant source IDs, and actions already completed. Keep the underlying transcript or source records according to your retention policy so important details remain inspectable.
Test a corrected preference and an abandoned plan. A summary that presents an old suggestion as the current decision is a memory error even if its prose reads smoothly.
Add context across conversations
A separate store namespace can serve multiple threads. The LangGraph memory guide shows that native distinction. A managed service can add retrieval, extraction, profiles or lifecycle capabilities alongside that native persistence.
The Supermemory LangChain integration documents the external memory path. Resolve identity on the server, retrieve a bounded amount of user context before generation, and store permitted source content afterward. Keep provider writes observable and distinguish accepted ingestion from searchable readiness.
A thread ID should identify the conversation. A user-memory scope should identify the authorized user within a tenant. Using the same global scope for every conversation makes cross-user contamination possible; using a random scope on every request prevents continuity.
Validate storage and answers separately
Use one test set for framework behavior: same-thread accumulation, different-thread separation, durable restart, and valid tool-message ordering. Use another for retrieval and generation: relevant preference available, irrelevant history ignored, corrections honored, and deleted material absent.
Run both sets before deployment. Record the package versions, model, prompts and corpus with the results so you can reproduce a failure after an upgrade.
The example uses Python 3.12, langchain 1.4.1 and langgraph 1.2.11.
When your LangChain agent needs context beyond its current thread, start a Supermemory integration using the documented path above. Keep native conversation history and compare what the additional retrieval contributes across two sessions.