Mental model: what a chunk actually is
A chunk is the unit your embedding model encodes into a vector. Every retrieval query compares against those vectors, so the chunk's content directly determines what the LLM sees. Think of each chunk as a candidate answer to some unknown future question. A well-formed chunk is semantically self-contained -- a reader dropped into it mid-document should understand the idea without needing surrounding paragraphs. A poorly-formed chunk is half a sentence, or three loosely-related paragraphs, or a table header with no rows. Embedding models compress a chunk into a fixed-size vector; the more semantic noise in the chunk, the less precise that vector is.
Overlap deserves its own paragraph. When you split at fixed boundaries you inevitably cut across sentences. A 50-token overlap means each chunk starts 50 tokens into the previous chunk's territory. This raises retrieval recall because a query that aligns with the end of chunk N can now match the beginning of chunk N+1. But overlap inflates your vector store by roughly overlap/chunk_size percent and means you retrieve near-duplicate content. Track it as a cost, not a free win.
Fixed-size chunking: how it works and when it breaks
The implementation is trivial: tokenize the document, slice at every N tokens, step forward by (N minus overlap) tokens. LangChain's CharacterTextSplitter and TokenTextSplitter, as well as LlamaIndex's SentenceSplitter with a fixed size, all do this. The predictability is the main value: every chunk has a known maximum token count, so you never accidentally exceed your embedding model's context limit (512 tokens for most SBERT models, 8191 for text-embedding-3-large).
Fixed-size breaks badly on structured documents: legal contracts split mid-clause, code examples split mid-function, tables split mid-row. It also ignores that information density is uneven -- a dense technical paragraph might pack more meaning into 100 tokens than a transitional paragraph packs into 500. Use fixed-size as a baseline to beat, not as a default to ship.
Recursive chunking: the practical default
Recursive chunking tries a hierarchy of separators -- typically ["\n\n", "\n", ". ", " ", ""] -- and only descends to the next separator if the current chunk still exceeds the target size. This means paragraphs stay intact when they fit, sentences stay intact when paragraphs are too long, and only as a last resort does it split mid-word. LangChain's RecursiveCharacterTextSplitter is the canonical implementation.
For most prose (blog posts, documentation, research papers, support tickets), recursive chunking produces meaningfully better retrieval than fixed-size with almost no extra cost. The tradeoff is slightly variable chunk sizes, which is fine for dense vector retrieval but worth knowing if downstream code assumes uniform chunk lengths. For code, swap the separator list to ["\nclass ", "\ndef ", "\n\n", "\n", " ", ""] to keep function definitions intact.
Semantic chunking: when meaning overrides structure
Semantic chunking computes sentence-level embeddings, calculates the cosine distance between consecutive sentences, and inserts a chunk boundary wherever the distance exceeds a threshold. The intuition: when adjacent sentences are semantically distant, a topic shift has occurred and a new chunk should start.
The practical implementation (available in LlamaIndex as SemanticSplitterNodeParser) embeds every sentence individually, computes a rolling similarity window, and finds breakpoints at local minima of similarity. The cost is real: for a 10,000-word document with ~500 sentences you pay ~500 embedding API calls (or one batch call) at ingest time. At $0.02 per million tokens for text-embedding-3-small, that is negligible per document, but adds up when re-ingesting a large corpus. More importantly, this makes ingest time dependent on embedding API latency, which matters for real-time ingestion pipelines.
Semantic chunking wins clearly when documents contain multiple distinct topics in sequence -- a Wikipedia article that covers background, technical detail, controversy, and reception in sequence will produce far better chunks than any structural splitter. It is harder to justify for already-structured documents (API reference docs, FAQs) where the structure already maps to semantics.
A real scenario: internal knowledge base for a SaaS product
You have 800 Markdown files: release notes, how-to guides, and API reference docs. A user asks "how do I configure SSO with Okta?" The relevant content is in three places: a conceptual overview paragraph, a step-by-step guide section, and a config object reference. Here is how each strategy handles ingest:
Fixed-size with 512 tokens and 50-token overlap: the step-by-step guide probably fits in 2-3 chunks cleanly. The config object table might split mid-row. Retrieval works but the LLM sometimes gets half a table.
Recursive: the \n\n separator keeps each numbered step as its own chunk when steps are short. The table stays intact if it fits in 512 tokens. Better.
Semantic: the conceptual overview, step-by-step, and config reference get their own chunks naturally because their embedding distances are high. At retrieval time, all three surface cleanly. Best retrieval quality but 800-file ingest now makes ~40,000 embedding calls (50 sentences/file average).
A senior engineer would use recursive chunking for the first production version, instrument retrieval recall with a labeled eval set of 50 real user questions, and only invest in semantic chunking for the document types where recursive chunking demonstrably underperforms.
What changes at scale
At 10 users, chunking strategy barely matters -- your eval corpus is small, latency is fine, cost is negligible. At 10,000 users you start seeing patterns: users ask about content that spans chunk boundaries, certain document types retrieve poorly, some chunks appear in nearly every result because they contain generic preamble text ("This document describes..."). You add chunk-level metadata filtering and start tuning overlap.
At 10 million documents, the cost of semantic chunking at ingest is non-trivial, and you want incremental re-chunking when documents update. You also start caring about chunk deduplication -- near-identical chunks from versioned documents pollute your vector store and inflate retrieval cost. At that scale, chunking is part of your data pipeline SLA, not a one-time script.
Key Takeaways
- Fixed-size chunking is fast but splits semantic units; always add overlap to reduce boundary damage.
- Recursive chunking respects document structure; prefer it over fixed-size for most prose corpora.
- Semantic chunking maximizes retrieval relevance but adds embedding cost and latency at ingest time.
- Measure chunk quality empirically using retrieval recall on a labeled eval set, not intuition.
Pro tips
- Measure chunking quality with retrieval recall on a labeled eval set before shipping. Pick 40-60 real user questions, label which document passages contain the answer, then check what percentage of top-5 retrieved chunks contain at least one labeled passage. This single metric tells you more than any intuition about chunk size.
- Chunk size and embedding model context length are coupled. Most SBERT-family open-source models cap at 512 tokens. If you set chunk_size=800 tokens, the embedding model silently truncates. Check the model card and set chunk_size to roughly 80% of the model's max token limit.
- Store chunk metadata (source file, section heading, chunk index, character offsets) at ingest time. You will need it later for citation rendering, deduplication, incremental re-ingestion, and debugging retrieval failures. Retrofitting metadata into an existing vector store is painful.
- Preamble chunks (intro sentences like 'This guide covers...') match many queries but contain no actionable content. Filter them out by length (drop chunks under ~80 tokens) or by running a one-time classifier. A handful of low-information high-frequency chunks can dominate your retrieval results and confuse the LLM.
Common pitfalls
- Mistake: Using character count for chunk_size when your embedding model has a token limit. Fix: Use a token-aware splitter like
TokenTextSplitteror set character limits conservatively (1 token ~ 4 chars for English). - Mistake: Setting overlap to zero. Fix: Use 10-15% overlap of chunk size. Boundary sentences carry meaning and queries often align with them.
- Mistake: Re-chunking the entire corpus on every document update. Fix: Track a content hash per document and only re-chunk and re-embed documents whose hash changed.
- Mistake: Applying the same chunking config to every document type. Fix: Use recursive chunking with code-specific separators for code files and prose-specific separators for articles; one config rarely fits both.
When to use fixed-size vs recursive vs semantic chunking
| Option | Use when | Avoid when |
|---|---|---|
| Fixed-size | You need a fast baseline, documents are uniform in structure, or you are prototyping and want predictable chunk counts. | Documents contain tables, code, numbered lists, or multi-topic prose where structural splits matter. |
| Recursive | Most production prose corpora: documentation, articles, support tickets. Balances quality and simplicity with near-zero extra cost. | Documents have no meaningful delimiters (raw OCR output, continuous stream data) or you need guaranteed chunk size uniformity. |
| Semantic | Documents cover multiple distinct topics in sequence (encyclopedic articles, long reports) and retrieval recall on recursive chunking is measurably low. | Ingest latency is constrained, you are re-ingesting frequently, or documents are already well-structured with clear section headers. |
| Hybrid (recursive + semantic) | Large heterogeneous corpora where some document types benefit from structure-awareness and others from semantic boundaries. | Your team does not have an eval pipeline to validate the added complexity is paying off in retrieval metrics. |
Code Example
# langchain-text-splitters==0.2.0
from langchain_text_splitters import RecursiveCharacterTextSplitter
document = """
## Authentication
All API requests require a Bearer token in the Authorization header.
Tokens expire after 24 hours.
## Rate Limits
Free tier: 60 requests/minute.
Pro tier: 600 requests/minute.
Exceeding the limit returns HTTP 429.
"""
splitter = RecursiveCharacterTextSplitter(
chunk_size=120, # characters, not tokens
chunk_overlap=20,
separators=["\n\n", "\n", ". ", " ", ""],
)
chunks = splitter.split_text(document)
for i, chunk in enumerate(chunks):
print(f"--- Chunk {i} ({len(chunk)} chars) ---")
print(chunk)How this code works
This code demonstrates a recursive text splitting strategy, which is vital for breaking down a larger document into smaller, digestible chunks suitable for Retrieval-Augmented Generation (RAG) pipelines. Its job is to ensure that a long input text is divided into segments that can fit within an LLM's context window while trying to maintain meaningful context within each piece.
The process begins by importing RecursiveCharacterTextSplitter and defining the multi-paragraph document to be processed. An instance of splitter is then created, configured with a chunk_size=120 and chunk_overlap=20. A key point for beginners is that chunk_size here refers to the number of characters, not tokens, which can be a source of confusion. The separators list is a critical feature, specifying an ordered set of delimiters (like "\n\n" for paragraphs, then "\n" for lines) that the splitter will try, in sequence, to find natural breaks. It attempts the first separator; if chunks are still too large, it recursively applies the next separator on those oversized chunks. Finally, splitter.split_text(document) executes this intelligent division, and the subsequent loop prints each resulting chunk along with its character count.
Production-grade example
Adds retries with backoff, token cost logging, batching, and per-error structured logs vs the basic example.
# langchain-text-splitters==0.2.0, openai==1.30.0, tenacity==8.3.0
import logging
import os
import time
from typing import Iterator
from langchain_text_splitters import RecursiveCharacterTextSplitter
from openai import OpenAI, RateLimitError, APITimeoutError
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
logger = logging.getLogger(__name__)
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=30.0)
SPLITTER = RecursiveCharacterTextSplitter(
chunk_size=400,
chunk_overlap=40,
separators=["\n\n", "\n", ". ", " ", ""],
)
@retry(
retry=retry_if_exception_type((RateLimitError, APITimeoutError)),
wait=wait_exponential(multiplier=1, min=2, max=60),
stop=stop_after_attempt(5),
)
def embed_batch(texts: list[str]) -> list[list[float]]:
response = client.embeddings.create(
model="text-embedding-3-small",
input=texts,
)
total_tokens = response.usage.total_tokens
cost_usd = total_tokens / 1_000_000 * 0.02 # illustrative; verify current pricing
logger.info("embed_batch size=%d tokens=%d est_cost_usd=%.6f", len(texts), total_tokens, cost_usd)
return [item.embedding for item in response.data]
def chunk_and_embed(
doc_id: str,
text: str,
batch_size: int = 64,
) -> Iterator[dict]:
chunks = SPLITTER.split_text(text)
if not chunks:
logger.warning("doc_id=%s produced zero chunks; skipping", doc_id)
return
logger.info("doc_id=%s chunk_count=%d", doc_id, len(chunks))
for batch_start in range(0, len(chunks), batch_size):
batch = chunks[batch_start : batch_start + batch_size]
try:
vectors = embed_batch(batch)
except Exception as exc:
logger.error("doc_id=%s batch_start=%d embed_failed error=%r", doc_id, batch_start, exc)
raise # let the caller decide whether to skip or abort
for offset, (chunk_text, vector) in enumerate(zip(batch, vectors)):
yield {
"id": f"{doc_id}#{batch_start + offset}",
"doc_id": doc_id,
"chunk_index": batch_start + offset,
"text": chunk_text,
"embedding": vector,
}
time.sleep(0.05) # gentle rate-limit buffer between batchesHow this code works
This Python code's job is to prepare large text documents for a Retrieval-Augmented Generation (RAG) system. It efficiently breaks down lengthy texts into smaller, semantically meaningful pieces called "chunks" and then transforms these chunks into numerical representations known as "embeddings." These embeddings are crucial for AI models to quickly understand and retrieve relevant information from a vast document collection. The primary orchestrator is the chunk_and_embed function, which takes a document ID and its raw text, then systematically processes it for embedding.
The process begins with the SPLITTER, a RecursiveCharacterTextSplitter, which intelligently divides the input text into chunks, attempting to break at logical points defined by separators like newlines, while adhering to a chunk_size and chunk_overlap. Next, the embed_batch function sends these chunks to the OpenAI API to generate embeddings. A key feature here is the @retry decorator using tenacity, which automatically re-attempts API calls if it encounters common issues like RateLimitError or APITimeoutError, ensuring resilience against temporary service interruptions. The chunk_and_embed function then batches these chunks, calls embed_batch, and yields structured dictionaries, each containing a chunk's text, its embedding, and identifying metadata. A subtle but important detail is the time.sleep(0.05) after each batch, which acts as a gentle buffer to prevent overwhelming the API, even before the tenacity retry mechanism might activate.
Practice & master
Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.
Exercise
Build a small chunking comparison script. Given a 500-word sample text (provided in starter code), chunk it with both RecursiveCharacterTextSplitter and a naive fixed-size splitter. For each strategy, print the number of chunks, the min/max/average chunk length in characters, and the first 80 characters of each chunk. Experiment with chunk_size=200 and chunk_size=400.
# langchain-text-splitters==0.2.0
from langchain_text_splitters import RecursiveCharacterTextSplitter, CharacterTextSplitter
SAMPLE = """
Vector databases store high-dimensional embeddings and support approximate nearest-neighbor search.
They are a core component of modern RAG pipelines.
Pinecone is a managed service that handles indexing and scaling automatically.
Weaviate is open-source and supports hybrid BM25 plus vector search out of the box.
Qdrant offers a Rust-based engine with low memory overhead.
Choosing the right vector database depends on your latency requirements, data volume,
and whether you need hybrid search. For most teams starting out, a hosted option
reduces operational burden significantly. As data grows past tens of millions of vectors,
cost and query latency become the primary drivers of the decision.
"""
def analyze_chunks(name: str, chunks: list[str]) -> None:
# TODO: print count, min length, max length, average length
# TODO: print first 80 chars of each chunk
pass
for size in [200, 400]:
# TODO: create RecursiveCharacterTextSplitter with chunk_size=size, chunk_overlap=20
# TODO: create CharacterTextSplitter with chunk_size=size, chunk_overlap=20, separator=" "
# TODO: split SAMPLE with each splitter and call analyze_chunks
passQuick check
A document has clear section headings and numbered steps. Which chunking strategy preserves that structure most reliably?
Why does chunk overlap increase retrieval recall rather than precision?
You add semantic chunking and ingest time triples. The most likely root cause is: