At the lowest level, every agent run is a sequence of LLM calls. Between calls, the agent needs to carry state. That state is memory. The problem is that LLM context windows are finite (and expensive per token), disk and database lookups add latency, and naive approaches — like shoving everything into the system prompt — degrade performance as the prompt grows. Designing memory well means choosing the right store for the right data at the right time.
Short-term memory (STM) is what lives in the active context window during a single agent run. It includes the current conversation history, tool outputs from this session, and scratchpad reasoning (like the Thought/Observation pairs from a ReAct loop). The implementation is usually a rolling buffer: keep the last N messages, or summarize older ones and replace them. LangChain's ConversationBufferWindowMemory and ConversationSummaryBufferMemory are both implementations of this pattern. The key constraint is token budget: on gpt-4o with a 128k context, you have ~96k tokens to work with after system prompt, tools, and response reservation. At 10 messages per minute, a long session will blow that ceiling in under two hours. The pro move is to set a hard token limit on the STM buffer and trigger summarization automatically when you approach it — not reactively after hitting a 400 error.
Long-term memory (LTM) is persistent across sessions and agent runs. It stores user preferences, domain facts, learned procedures, and anything that should survive a server restart. The canonical implementation is a vector database (Pinecone, Qdrant, pgvector) where facts are encoded as embeddings and retrieved by semantic similarity at query time. The agent injects the top-k retrieved facts into the system prompt before each LLM call. This is essentially RAG applied to agent state. The tricky part is write policy: when do you decide a new fact is worth storing? A simple heuristic is to write to LTM at the end of each session only if a user explicitly confirmed or corrected something. Blind writes from every agent observation pollute the store quickly and lead to contradictory facts with no resolution mechanism.
Episodic memory is a timestamped log of specific events: "On 2024-11-15, user Priya ran a data pipeline job that failed because the S3 bucket credentials expired. She resolved it by rotating the AWS key." That record is fundamentally different from the LTM fact "Priya uses AWS S3." Episodic memories encode the narrative arc of what happened, not just the static outcome. Retrieval is typically hybrid: you pull episodes by recency (last 5 sessions), by relevance (embed the current task and find similar past episodes), or by explicit tag ("failures related to AWS"). A practical pattern is to store episodes in a structured format — a JSON object with fields like timestamp, task, outcome, error_type, resolution — and index the task field as an embedding. This lets you query: "what did we do last time we hit this error?"
The interplay between all three during a live agent run looks like this: at session start, the agent loads relevant LTM facts and the last 3 episodic summaries into the system prompt. During the run, new observations accumulate in STM. At session end, the agent runs a summarization step: important new facts are written to LTM, and the session is logged as a new episode. This write-at-end pattern keeps the hot path (inference) fast while ensuring durability.
At scale, the failure modes shift. With 10 users you can store everything and retrieve it all. At 10k users, LTM grows large enough that embedding search latency matters — you need proper ANN indexing and namespaced retrieval so user A's memory doesn't bleed into user B's context. At 10M users, the memory system becomes a primary cost center. A single user's LTM read per session is cheap; 10M concurrent sessions means millions of vector queries per minute. You'll need caching (Redis for hot user memory), tiered storage (recent memories in Qdrant, archived in S3 with lazy load), and memory eviction policies. Episodic logs at this scale also require a retention policy — storing every session forever is not sustainable. A rolling 90-day window with monthly summaries is a common production pattern.
Key Takeaways
- Short-term memory lives in the context window; manage it actively or you'll hit token limits.
- Long-term memory needs semantic retrieval — a plain list doesn't scale past a few hundred facts.
- Episodic memory enables personalization and debugging; store structured events, not raw text.
- Never mix memory types in the same store; retrieval semantics and TTLs differ fundamentally.
Pro tips
- Set a
score_thresholdon every LTM retrieval call. Without it, you'll inject the 'least bad' facts even when none are relevant, which causes the agent to confidently confabulate details from weakly-related past sessions. - Summarize STM before it overflows, not after. Use a token-counting utility (tiktoken) to monitor buffer size on every write, and trigger a summarization chain when you cross 70% of your reserved context budget.
- Use separate Qdrant collections (or namespaces in Pinecone) per memory type. Mixing episodic logs and LTM facts in one collection means your retrieval scores are incomparable — a high-similarity episode will outrank a perfectly relevant fact.
- Write to LTM at session end, not mid-session. Mid-session writes create race conditions if the agent contradicts itself within a single run, and you get duplicate or conflicting facts. A single end-of-session consolidation step is cleaner.
Common pitfalls
- Mistake: Injecting all retrieved LTM facts into every LLM call unconditionally. Fix: Filter by
score_thresholdand cap at 5 facts; more facts degrade reasoning quality and inflate costs. - Mistake: Storing raw conversation turns as episodic memory. Fix: Summarize each session into a structured JSON record with explicit fields (task, outcome, error_type) before writing; raw turns are hard to retrieve semantically.
- Mistake: Sharing one memory namespace across all users. Fix: Always filter by
user_idin your vector DB query; cross-user memory leakage is a privacy incident waiting to happen. - Mistake: Never evicting old LTM facts, letting contradictions accumulate. Fix: During writes, search for near-duplicate facts and delete or merge them before inserting the new one.
Which memory type to use for a given piece of agent state
| Option | Use when | Avoid when |
|---|---|---|
| Short-term (in-memory buffer) | Data is only relevant for the current task or session; context window fits it; latency must be sub-millisecond. | Data needs to survive session end or be shared across agent instances; buffer would exceed token budget. |
| Long-term (vector DB) | Storing persistent facts, user preferences, or domain knowledge that must be semantically retrievable across sessions. | Data is highly time-sensitive or narrative (what happened when); retrieval by recency matters more than similarity. |
| Episodic (structured event log + embeddings) | You need to recall what happened in a prior session, learn from past successes or failures, or personalize based on history. | Session volume is massive and you haven't implemented a retention/eviction policy; storage costs will balloon. |
| Summarized STM (ConversationSummaryBufferMemory) | Sessions are long enough to risk hitting context limits but you still need continuity within the session. | Summarization itself is too slow or expensive for your latency budget; raw buffer fits easily in context. |
Code Example
# langchain-core==0.2.x, langchain-openai==0.1.x
from langchain_openai import ChatOpenAI
from langchain.memory import ConversationBufferWindowMemory
from langchain.chains import ConversationChain
llm = ChatOpenAI(model="gpt-4o-mini", temperature=0)
# Short-term memory: keep only the last 5 exchanges (k=5)
memory = ConversationBufferWindowMemory(k=5, return_messages=True)
chain = ConversationChain(llm=llm, memory=memory, verbose=False)
response1 = chain.predict(input="My name is Priya and I prefer Python over JS.")
response2 = chain.predict(input="What language do I prefer?")
print(response2) # Should recall Python preference from the buffer
print(f"Messages in buffer: {len(memory.chat_memory.messages)}")How this code works
This code example demonstrates how to implement short-term memory for an AI agent using LangChain. Its job is to allow the AI to remember recent pieces of information within a conversation, such as a user's stated preference, and successfully recall it a few turns later. The ChatOpenAI component acts as the language model or the agent's "brain." The core of the short-term recall here is ConversationBufferWindowMemory, which stores a limited number of past conversation turns. In this setup, k=5 means it will only remember the last five exchanges, creating a rolling window of the most recent dialogue.
The ConversationChain then combines the llm (brain) and this specific memory buffer. When chain.predict processes inputs like "My name is Priya and I prefer Python over JS." followed by "What language do I prefer?", the agent successfully answers "Python" because the preference is still within the k=5 window of memory. A subtle but important detail is that return_messages=True is specified; while often the default, it ensures the memory buffer stores actual message objects, which is standard practice for many LangChain chains. If the conversation had exceeded five turns between the preference statement and the question, the agent would "forget" the initial preference, highlighting the transient nature of short-term, windowed memory. The final len(memory.chat_memory.messages) confirms the current number of messages held in this limited buffer.
Production-grade example
Adds retries with backoff, token logging, score-threshold filtering, structured episode writes, and env-var auth.
# langchain-openai==0.1.x, langchain-community==0.2.x, openai==1.x, qdrant-client==1.9.x
import os
import logging
import time
from datetime import datetime
from typing import Optional
from openai import OpenAI, RateLimitError, APITimeoutError
from qdrant_client import QdrantClient
from qdrant_client.models import PointStruct, Distance, VectorParams
import uuid
logger = logging.getLogger(__name__)
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=15.0)
qdrant = QdrantClient(url=os.environ["QDRANT_URL"], api_key=os.environ["QDRANT_API_KEY"])
COLLECTION = "agent_ltm"
EMBED_MODEL = "text-embedding-3-small"
def ensure_collection():
existing = [c.name for c in qdrant.get_collections().collections]
if COLLECTION not in existing:
qdrant.create_collection(COLLECTION, vectors_config=VectorParams(size=1536, distance=Distance.COSINE))
def embed_with_retry(text: str, max_attempts: int = 3) -> list[float]:
for attempt in range(max_attempts):
try:
resp = client.embeddings.create(input=text, model=EMBED_MODEL)
tokens_used = resp.usage.total_tokens
logger.info("embed tokens=%d attempt=%d", tokens_used, attempt + 1)
return resp.data[0].embedding
except RateLimitError:
wait = 2 ** attempt
logger.warning("rate limit hit, sleeping %ds", wait)
time.sleep(wait)
except APITimeoutError:
logger.error("embed timeout on attempt %d", attempt + 1)
if attempt == max_attempts - 1:
raise
raise RuntimeError("embed_with_retry exhausted attempts")
def write_ltm_fact(user_id: str, fact: str) -> str:
ensure_collection()
vector = embed_with_retry(fact)
point_id = str(uuid.uuid4())
qdrant.upsert(COLLECTION, points=[PointStruct(
id=point_id,
vector=vector,
payload={"user_id": user_id, "fact": fact, "written_at": datetime.utcnow().isoformat()}
)])
logger.info("ltm_write user=%s point=%s", user_id, point_id)
return point_id
def retrieve_ltm_facts(user_id: str, query: str, top_k: int = 5) -> list[str]:
vector = embed_with_retry(query)
results = qdrant.search(
collection_name=COLLECTION,
query_vector=vector,
query_filter={"must": [{"key": "user_id", "match": {"value": user_id}}]},
limit=top_k,
score_threshold=0.75 # avoid injecting low-relevance facts
)
facts = [r.payload["fact"] for r in results]
logger.info("ltm_retrieve user=%s query=%r hits=%d", user_id, query[:60], len(facts))
return facts
def write_episode(user_id: str, task: str, outcome: str, error_type: Optional[str] = None):
episode = {"timestamp": datetime.utcnow().isoformat(), "task": task, "outcome": outcome, "error_type": error_type}
# In production: append to a structured store (Postgres, BigQuery, etc.)
logger.info("episode_write user=%s episode=%s", user_id, episode)
return episodeHow this code works
This Python code establishes memory systems for an AI agent, specifically handling long-term and episodic memory. It utilizes OpenAI to transform text into numerical embeddings (vectors) and Qdrant, a vector database, to store and retrieve these memories. The OpenAI and qdrant clients are initialized using environment variables for secure access, and ensure_collection sets up the necessary storage within Qdrant, preparing it to hold agent facts. The system allows agents to persistently store information and intelligently recall relevant details.
For long-term memory, the embed_with_retry function is crucial; it converts text facts or query strings into vectors. This function notably includes a robust retry mechanism with exponential time.sleep to gracefully manage common RateLimitError or APITimeoutError from the embedding service, a pattern often overlooked by beginners that can prevent frustrating program crashes. The write_ltm_fact function stores these embedded facts with a user_id and timestamp, while retrieve_ltm_facts searches Qdrant for relevant memories, filtering by user_id and applying a score_threshold (e.g., 0.75) to ensure only highly pertinent information is returned. Episodic memory, managed by write_episode, logs agent interactions like task and outcome, serving as a structured record of events.
Practice & master
Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.
Exercise
Build a simple agent memory manager that: (1) stores up to 5 messages in a short-term buffer, (2) summarizes the buffer when it hits 5 messages and writes the summary as a long-term fact to a local in-memory list, and (3) retrieves the 3 most recent LTM facts and prepends them to any new conversation context. Use the OpenAI chat completions API directly (no LangChain).
import os
from openai import OpenAI
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
STM_LIMIT = 5
stm_buffer: list[dict] = [] # list of {"role": ..., "content": ...}
ltm_store: list[str] = [] # list of summary strings
def summarize_buffer(messages: list[dict]) -> str:
# TODO: call client.chat.completions.create with a summarization prompt
# Return a single string summarizing the key facts from messages
pass
def add_message(role: str, content: str):
# TODO: append to stm_buffer
# If len(stm_buffer) >= STM_LIMIT, call summarize_buffer,
# append the result to ltm_store, then clear stm_buffer
pass
def build_context(user_query: str) -> list[dict]:
# TODO: take the last 3 items from ltm_store, format them as a system message,
# then append stm_buffer and the new user_query
# Return a messages list ready for chat.completions.create
pass
if __name__ == "__main__":
for turn in ["My name is Alex.", "I work on distributed systems.",
"I prefer Go over Python.", "I dislike YAML configs.",
"I love observability tooling.", "What do you know about me?"]:
add_message("user", turn)
context = build_context(turn)
resp = client.chat.completions.create(model="gpt-4o-mini", messages=context)
reply = resp.choices[0].message.content
add_message("assistant", reply)
print(f"User: {turn}\nAgent: {reply}\n")Quick check
An agent is mid-session and its STM buffer is approaching the context limit. What is the correct next step?
You store a user's LTM facts without a user_id filter on retrieval. What is the most likely production failure?
What fundamentally distinguishes episodic memory from long-term memory in an agent system?