Phase 3: RAG & Knowledge Systems

Chunking strategies: fixed-size, semantic & recursive

Advanced ~16 min read
Think of it this way A friendly analogy. Read this if the technical version feels dense. Show Hide

Imagine you have a super smart robot chef that knows how to make anything, but it learns from all sorts of recipe books, cooking blogs, and family recipe cards. The challenge is, how do you give it all that information so it can actually find what it needs quickly? If you give it a whole giant cookbook as one single, massive scroll, it's like trying to find "how to make chocolate chip cookies" in a book with no pages, no chapters, and no table of contents! It would take forever and might miss something important.

On the other hand, if you cut every single word like "flour" or "mix" onto its own tiny little slip of paper, your robot chef would get lost too. It would have a million slips but no idea how they fit together to make a whole recipe. So, you need to break down all that cooking knowledge into "just right" sized pieces. This is what we call "chunking" – deciding how to slice up your big piles of information so your robot can use them effectively.

There are different clever ways to do this. One way is fixed-size: you just decide every piece of information must be exactly 10 lines long, no matter what. Sometimes this works great, but sometimes it cuts a recipe for cookies right in half, like "add 2 cups of..." on one piece, and "flour" on the next! Another way is recursive: this is smarter. It tries to cut by big, natural breaks first, like "chapters" (all the dessert recipes). If a chapter is still too big, it then tries to cut by smaller breaks, like individual "recipes." It keeps trying smaller and smaller, more sensible cuts until each piece is a good size without chopping up important steps.

Then there's the really smart way called semantic chunking. This method tries to understand what each part of the recipe is about. It's like your robot chef reads the whole book and says, "Aha! All these sentences are talking about baking time, and all these sentences are about ingredients." It groups pieces of information together based on what they mean, even if they are far apart in the original book. This means when you build tools that help computers learn from tons of text, you get to choose how it best organizes all that knowledge so it can find the right answers fast, helping it bake the perfect virtual cookies every time!

Mental model: what a chunk actually is

A chunk is the unit your embedding model encodes into a vector. Every retrieval query compares against those vectors, so the chunk's content directly determines what the LLM sees. Think of each chunk as a candidate answer to some unknown future question. A well-formed chunk is semantically self-contained -- a reader dropped into it mid-document should understand the idea without needing surrounding paragraphs. A poorly-formed chunk is half a sentence, or three loosely-related paragraphs, or a table header with no rows. Embedding models compress a chunk into a fixed-size vector; the more semantic noise in the chunk, the less precise that vector is.

Overlap deserves its own paragraph. When you split at fixed boundaries you inevitably cut across sentences. A 50-token overlap means each chunk starts 50 tokens into the previous chunk's territory. This raises retrieval recall because a query that aligns with the end of chunk N can now match the beginning of chunk N+1. But overlap inflates your vector store by roughly overlap/chunk_size percent and means you retrieve near-duplicate content. Track it as a cost, not a free win.

Fixed-size chunking: how it works and when it breaks

The implementation is trivial: tokenize the document, slice at every N tokens, step forward by (N minus overlap) tokens. LangChain's CharacterTextSplitter and TokenTextSplitter, as well as LlamaIndex's SentenceSplitter with a fixed size, all do this. The predictability is the main value: every chunk has a known maximum token count, so you never accidentally exceed your embedding model's context limit (512 tokens for most SBERT models, 8191 for text-embedding-3-large).

Fixed-size breaks badly on structured documents: legal contracts split mid-clause, code examples split mid-function, tables split mid-row. It also ignores that information density is uneven -- a dense technical paragraph might pack more meaning into 100 tokens than a transitional paragraph packs into 500. Use fixed-size as a baseline to beat, not as a default to ship.

Recursive chunking: the practical default

Recursive chunking tries a hierarchy of separators -- typically ["\n\n", "\n", ". ", " ", ""] -- and only descends to the next separator if the current chunk still exceeds the target size. This means paragraphs stay intact when they fit, sentences stay intact when paragraphs are too long, and only as a last resort does it split mid-word. LangChain's RecursiveCharacterTextSplitter is the canonical implementation.

For most prose (blog posts, documentation, research papers, support tickets), recursive chunking produces meaningfully better retrieval than fixed-size with almost no extra cost. The tradeoff is slightly variable chunk sizes, which is fine for dense vector retrieval but worth knowing if downstream code assumes uniform chunk lengths. For code, swap the separator list to ["\nclass ", "\ndef ", "\n\n", "\n", " ", ""] to keep function definitions intact.

Semantic chunking: when meaning overrides structure

Semantic chunking computes sentence-level embeddings, calculates the cosine distance between consecutive sentences, and inserts a chunk boundary wherever the distance exceeds a threshold. The intuition: when adjacent sentences are semantically distant, a topic shift has occurred and a new chunk should start.

The practical implementation (available in LlamaIndex as SemanticSplitterNodeParser) embeds every sentence individually, computes a rolling similarity window, and finds breakpoints at local minima of similarity. The cost is real: for a 10,000-word document with ~500 sentences you pay ~500 embedding API calls (or one batch call) at ingest time. At $0.02 per million tokens for text-embedding-3-small, that is negligible per document, but adds up when re-ingesting a large corpus. More importantly, this makes ingest time dependent on embedding API latency, which matters for real-time ingestion pipelines.

Semantic chunking wins clearly when documents contain multiple distinct topics in sequence -- a Wikipedia article that covers background, technical detail, controversy, and reception in sequence will produce far better chunks than any structural splitter. It is harder to justify for already-structured documents (API reference docs, FAQs) where the structure already maps to semantics.

A real scenario: internal knowledge base for a SaaS product

You have 800 Markdown files: release notes, how-to guides, and API reference docs. A user asks "how do I configure SSO with Okta?" The relevant content is in three places: a conceptual overview paragraph, a step-by-step guide section, and a config object reference. Here is how each strategy handles ingest:

Fixed-size with 512 tokens and 50-token overlap: the step-by-step guide probably fits in 2-3 chunks cleanly. The config object table might split mid-row. Retrieval works but the LLM sometimes gets half a table.

Recursive: the \n\n separator keeps each numbered step as its own chunk when steps are short. The table stays intact if it fits in 512 tokens. Better.

Semantic: the conceptual overview, step-by-step, and config reference get their own chunks naturally because their embedding distances are high. At retrieval time, all three surface cleanly. Best retrieval quality but 800-file ingest now makes ~40,000 embedding calls (50 sentences/file average).

A senior engineer would use recursive chunking for the first production version, instrument retrieval recall with a labeled eval set of 50 real user questions, and only invest in semantic chunking for the document types where recursive chunking demonstrably underperforms.

What changes at scale

At 10 users, chunking strategy barely matters -- your eval corpus is small, latency is fine, cost is negligible. At 10,000 users you start seeing patterns: users ask about content that spans chunk boundaries, certain document types retrieve poorly, some chunks appear in nearly every result because they contain generic preamble text ("This document describes..."). You add chunk-level metadata filtering and start tuning overlap.

At 10 million documents, the cost of semantic chunking at ingest is non-trivial, and you want incremental re-chunking when documents update. You also start caring about chunk deduplication -- near-identical chunks from versioned documents pollute your vector store and inflate retrieval cost. At that scale, chunking is part of your data pipeline SLA, not a one-time script.

Key Takeaways

  • Fixed-size chunking is fast but splits semantic units; always add overlap to reduce boundary damage.
  • Recursive chunking respects document structure; prefer it over fixed-size for most prose corpora.
  • Semantic chunking maximizes retrieval relevance but adds embedding cost and latency at ingest time.
  • Measure chunk quality empirically using retrieval recall on a labeled eval set, not intuition.

Pro tips

  • Measure chunking quality with retrieval recall on a labeled eval set before shipping. Pick 40-60 real user questions, label which document passages contain the answer, then check what percentage of top-5 retrieved chunks contain at least one labeled passage. This single metric tells you more than any intuition about chunk size.
  • Chunk size and embedding model context length are coupled. Most SBERT-family open-source models cap at 512 tokens. If you set chunk_size=800 tokens, the embedding model silently truncates. Check the model card and set chunk_size to roughly 80% of the model's max token limit.
  • Store chunk metadata (source file, section heading, chunk index, character offsets) at ingest time. You will need it later for citation rendering, deduplication, incremental re-ingestion, and debugging retrieval failures. Retrofitting metadata into an existing vector store is painful.
  • Preamble chunks (intro sentences like 'This guide covers...') match many queries but contain no actionable content. Filter them out by length (drop chunks under ~80 tokens) or by running a one-time classifier. A handful of low-information high-frequency chunks can dominate your retrieval results and confuse the LLM.

Common pitfalls

  • Mistake: Using character count for chunk_size when your embedding model has a token limit. Fix: Use a token-aware splitter like TokenTextSplitter or set character limits conservatively (1 token ~ 4 chars for English).
  • Mistake: Setting overlap to zero. Fix: Use 10-15% overlap of chunk size. Boundary sentences carry meaning and queries often align with them.
  • Mistake: Re-chunking the entire corpus on every document update. Fix: Track a content hash per document and only re-chunk and re-embed documents whose hash changed.
  • Mistake: Applying the same chunking config to every document type. Fix: Use recursive chunking with code-specific separators for code files and prose-specific separators for articles; one config rarely fits both.

When to use fixed-size vs recursive vs semantic chunking

Option Use when Avoid when
Fixed-size You need a fast baseline, documents are uniform in structure, or you are prototyping and want predictable chunk counts. Documents contain tables, code, numbered lists, or multi-topic prose where structural splits matter.
Recursive Most production prose corpora: documentation, articles, support tickets. Balances quality and simplicity with near-zero extra cost. Documents have no meaningful delimiters (raw OCR output, continuous stream data) or you need guaranteed chunk size uniformity.
Semantic Documents cover multiple distinct topics in sequence (encyclopedic articles, long reports) and retrieval recall on recursive chunking is measurably low. Ingest latency is constrained, you are re-ingesting frequently, or documents are already well-structured with clear section headers.
Hybrid (recursive + semantic) Large heterogeneous corpora where some document types benefit from structure-awareness and others from semantic boundaries. Your team does not have an eval pipeline to validate the added complexity is paying off in retrieval metrics.

Code Example

python
# langchain-text-splitters==0.2.0
from langchain_text_splitters import RecursiveCharacterTextSplitter

document = """
## Authentication

All API requests require a Bearer token in the Authorization header.
Tokens expire after 24 hours.

## Rate Limits

Free tier: 60 requests/minute.
Pro tier: 600 requests/minute.
Exceeding the limit returns HTTP 429.
"""

splitter = RecursiveCharacterTextSplitter(
    chunk_size=120,        # characters, not tokens
    chunk_overlap=20,
    separators=["\n\n", "\n", ". ", " ", ""],
)

chunks = splitter.split_text(document)
for i, chunk in enumerate(chunks):
    print(f"--- Chunk {i} ({len(chunk)} chars) ---")
    print(chunk)

How this code works

This code demonstrates a recursive text splitting strategy, which is vital for breaking down a larger document into smaller, digestible chunks suitable for Retrieval-Augmented Generation (RAG) pipelines. Its job is to ensure that a long input text is divided into segments that can fit within an LLM's context window while trying to maintain meaningful context within each piece.

The process begins by importing RecursiveCharacterTextSplitter and defining the multi-paragraph document to be processed. An instance of splitter is then created, configured with a chunk_size=120 and chunk_overlap=20. A key point for beginners is that chunk_size here refers to the number of characters, not tokens, which can be a source of confusion. The separators list is a critical feature, specifying an ordered set of delimiters (like "\n\n" for paragraphs, then "\n" for lines) that the splitter will try, in sequence, to find natural breaks. It attempts the first separator; if chunks are still too large, it recursively applies the next separator on those oversized chunks. Finally, splitter.split_text(document) executes this intelligent division, and the subsequent loop prints each resulting chunk along with its character count.

Production-grade example

Adds retries with backoff, token cost logging, batching, and per-error structured logs vs the basic example.

python
# langchain-text-splitters==0.2.0, openai==1.30.0, tenacity==8.3.0
import logging
import os
import time
from typing import Iterator

from langchain_text_splitters import RecursiveCharacterTextSplitter
from openai import OpenAI, RateLimitError, APITimeoutError
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type

logger = logging.getLogger(__name__)
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")

client = OpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=30.0)

SPLITTER = RecursiveCharacterTextSplitter(
    chunk_size=400,
    chunk_overlap=40,
    separators=["\n\n", "\n", ". ", " ", ""],
)

@retry(
    retry=retry_if_exception_type((RateLimitError, APITimeoutError)),
    wait=wait_exponential(multiplier=1, min=2, max=60),
    stop=stop_after_attempt(5),
)
def embed_batch(texts: list[str]) -> list[list[float]]:
    response = client.embeddings.create(
        model="text-embedding-3-small",
        input=texts,
    )
    total_tokens = response.usage.total_tokens
    cost_usd = total_tokens / 1_000_000 * 0.02  # illustrative; verify current pricing
    logger.info("embed_batch size=%d tokens=%d est_cost_usd=%.6f", len(texts), total_tokens, cost_usd)
    return [item.embedding for item in response.data]

def chunk_and_embed(
    doc_id: str,
    text: str,
    batch_size: int = 64,
) -> Iterator[dict]:
    chunks = SPLITTER.split_text(text)
    if not chunks:
        logger.warning("doc_id=%s produced zero chunks; skipping", doc_id)
        return

    logger.info("doc_id=%s chunk_count=%d", doc_id, len(chunks))

    for batch_start in range(0, len(chunks), batch_size):
        batch = chunks[batch_start : batch_start + batch_size]
        try:
            vectors = embed_batch(batch)
        except Exception as exc:
            logger.error("doc_id=%s batch_start=%d embed_failed error=%r", doc_id, batch_start, exc)
            raise  # let the caller decide whether to skip or abort

        for offset, (chunk_text, vector) in enumerate(zip(batch, vectors)):
            yield {
                "id": f"{doc_id}#{batch_start + offset}",
                "doc_id": doc_id,
                "chunk_index": batch_start + offset,
                "text": chunk_text,
                "embedding": vector,
            }
        time.sleep(0.05)  # gentle rate-limit buffer between batches

How this code works

This Python code's job is to prepare large text documents for a Retrieval-Augmented Generation (RAG) system. It efficiently breaks down lengthy texts into smaller, semantically meaningful pieces called "chunks" and then transforms these chunks into numerical representations known as "embeddings." These embeddings are crucial for AI models to quickly understand and retrieve relevant information from a vast document collection. The primary orchestrator is the chunk_and_embed function, which takes a document ID and its raw text, then systematically processes it for embedding.

The process begins with the SPLITTER, a RecursiveCharacterTextSplitter, which intelligently divides the input text into chunks, attempting to break at logical points defined by separators like newlines, while adhering to a chunk_size and chunk_overlap. Next, the embed_batch function sends these chunks to the OpenAI API to generate embeddings. A key feature here is the @retry decorator using tenacity, which automatically re-attempts API calls if it encounters common issues like RateLimitError or APITimeoutError, ensuring resilience against temporary service interruptions. The chunk_and_embed function then batches these chunks, calls embed_batch, and yields structured dictionaries, each containing a chunk's text, its embedding, and identifying metadata. A subtle but important detail is the time.sleep(0.05) after each batch, which acts as a gentle buffer to prevent overwhelming the API, even before the tenacity retry mechanism might activate.

Practice & master

Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.

Exercise

Build a small chunking comparison script. Given a 500-word sample text (provided in starter code), chunk it with both RecursiveCharacterTextSplitter and a naive fixed-size splitter. For each strategy, print the number of chunks, the min/max/average chunk length in characters, and the first 80 characters of each chunk. Experiment with chunk_size=200 and chunk_size=400.

python
# langchain-text-splitters==0.2.0
from langchain_text_splitters import RecursiveCharacterTextSplitter, CharacterTextSplitter

SAMPLE = """
Vector databases store high-dimensional embeddings and support approximate nearest-neighbor search.
They are a core component of modern RAG pipelines.

Pinecone is a managed service that handles indexing and scaling automatically.
Weaviate is open-source and supports hybrid BM25 plus vector search out of the box.
Qdrant offers a Rust-based engine with low memory overhead.

Choosing the right vector database depends on your latency requirements, data volume,
and whether you need hybrid search. For most teams starting out, a hosted option
reduces operational burden significantly. As data grows past tens of millions of vectors,
cost and query latency become the primary drivers of the decision.
"""

def analyze_chunks(name: str, chunks: list[str]) -> None:
    # TODO: print count, min length, max length, average length
    # TODO: print first 80 chars of each chunk
    pass

for size in [200, 400]:
    # TODO: create RecursiveCharacterTextSplitter with chunk_size=size, chunk_overlap=20
    # TODO: create CharacterTextSplitter with chunk_size=size, chunk_overlap=20, separator=" "
    # TODO: split SAMPLE with each splitter and call analyze_chunks
    pass

Quick check

  1. A document has clear section headings and numbered steps. Which chunking strategy preserves that structure most reliably?

  2. Why does chunk overlap increase retrieval recall rather than precision?

  3. You add semantic chunking and ingest time triples. The most likely root cause is:

Self-check: Without referencing the lesson, explain to a colleague: when would you choose semantic chunking over recursive chunking, what metric would you use to confirm the choice was correct, and what operational cost does semantic chunking add at ingest time that recursive chunking does not?