Back

Retrieval-Augmented Generation in Production: Beyond the Demo

Feb 14, 2026 (6mo ago)

Retrieval-Augmented Generation in Production: Beyond the Demo

Every RAG demo works. You point a loader at a folder of PDFs, chop the text into 512-token blocks, embed them, drop them in a vector store, and ask three questions you already knew the answers to. It looks like magic.

Then you point it at 400,000 real documents — support tickets, contracts, a monorepo, a decade of Confluence — and it falls apart in ways that are genuinely hard to diagnose. This post is about the parts between the demo and the thing that survives contact with users.

The demo pipeline, and where it cracks

The naive pipeline is four steps: chunk → embed → cosine similarity top-k → stuff into the prompt.

Each step contains an assumption that real corpora violate.

  • Chunking assumes documents are prose. They aren't. They're tables, code, nested headings, invoices, log excerpts, and PDF two-column layouts that linearize into gibberish. A fixed-size splitter will cut a table away from its header row and a function away from its signature.
  • Embedding assumes semantic similarity is what you want. Often it isn't. If a user asks about error code ERR_CONN_RESET or part number MX-7741-B, embeddings will happily hand you passages about connection errors in general and never the one document that contains that literal string.
  • Top-k similarity assumes the retriever's ranking is the final ranking. A bi-encoder compresses a whole passage into one vector before it has ever seen your query. It's a coarse filter, not a judge.
  • Stuffing assumes the model reads everything you give it equally. It doesn't. Position matters, duplicates dilute, and contradictory chunks make the model hedge.

The fix isn't a better embedding model. It's turning one step into a pipeline of stages, each with a job it can actually do.

Chunking is the decision you regret longest

Chunking is upstream of everything. A bad chunk can never be retrieved correctly, no matter how good your reranker is — the information simply isn't in any single retrievable unit. And re-chunking means re-embedding, which means a full corpus rebuild.

Strategy How it works Good for Fails on
Fixed-size Split every N tokens Uniform prose, quick baselines Anything structured; cuts mid-sentence and mid-table
Recursive Split on a priority list of separators (paragraph, sentence, word) until chunks fit General documents Still blind to what the separators mean
Semantic Embed sentences, cut where consecutive similarity drops Long unstructured narrative Cost of embedding during ingestion; unstable on short docs
Structure-aware Parse the format first (headings, code AST, table rows), chunk along its natural boundaries Code, Markdown, HTML, spec docs, tables Needs a parser per format

My default is: structure-aware where a parser exists, recursive as the fallback, semantic only when I've proven the other two are losing. Semantic chunking is the one people reach for first because it sounds sophisticated, and it's the one that most often adds ingestion cost for a marginal retrieval gain.

Overlap is a patch, not a strategy

Overlap exists so that a fact straddling a boundary appears whole in at least one chunk. It works, and it's cheap insurance — a modest overlap on a recursive splitter is almost always worth it. But it inflates your index, it creates near-duplicate retrieval results you'll have to deduplicate later, and it does nothing for the real problem, which is that the boundary was in the wrong place. If you find yourself raising overlap to fix retrieval quality, the actual fix is upstream: chunk along structure instead.

Every chunk needs a header it didn't ask for

The single highest-leverage chunking trick is contextual prefixing: prepend the document's title and heading path to the chunk text before embedding it. A chunk that reads "Set the timeout to 30s." is nearly unretrievable. The same chunk embedded as "Billing API / Webhooks / Retry behaviour — Set the timeout to 30s." is retrievable by anyone asking about webhook retries.

That makes a chunk a record, not a string. What each field is carrying:

Field What it's for
text What the model sees
embed_text What you embed — the text with its structural context prepended
doc_id + ordinal Document and position within it, which together give a stable chunk_id of doc_id:ordinal
source_uri Where the chunk came from, for filtering and citation
heading_path ["Billing API", "Webhooks", "Retry behaviour"] — the prefix that goes into embed_text
kind "prose", "code", "table" or "list"
content_hash Diffing a re-ingested document against what is already indexed
meta Tenant, ACL, version, updated_at

The splitter that fills this in needs one rule the naive ones lack: a heading or the size limit is a cut point only outside a fenced block. Track whether you are inside one, and let the buffer run past max_chars rather than cut a fenced block in half or separate a table from its header row. Keep the heading path as a running stack, truncate it to the level of each new heading, and prepend it to embed_text as each chunk is flushed.

Attach whatever you'll ever want to filter or cite by: tenant, ACL group, document version, updated_at, source URI, page number, and the heading path. Metadata you didn't store at ingestion time is metadata you can only get back by re-indexing.

Embeddings: pick a model you can afford to keep

Three axes, and only one of them is quality.

Dimensionality. More dimensions means more bytes per vector, more work per comparison, and more index memory — and index memory is usually what your bill is actually made of. If your store supports it, dimension truncation on models trained for it (or product quantization) buys back a lot of that at a modest recall cost. Measure the recall cost on your corpus; it isn't a constant.

Domain fit. A general-purpose model trained on web text has seen very little that looks like your ICD-10 codes, your Verilog, or your Portuguese-language insurance policies. Domain fit beats leaderboard rank almost every time. Also check whether the model is asymmetric — trained with separate query and document prefixes. If it is and you skip the prefixes, you silently lose quality and nothing errors.

Switching cost. This is the one people underweight. Embeddings from different models are not comparable, so changing models means re-embedding the entire corpus, rebuilding the index, and re-running every retrieval evaluation you have. Budget for it from day one:

  • Store the raw chunk text and its metadata in a real database, not only in the vector store. The vector store is a derived index; you must be able to rebuild it from scratch without re-crawling sources.
  • Record the model name and version as a field on every vector. Mixed-model indexes are a silent correctness bug, not an error.
  • Make the embedding step idempotent and keyed on content_hash + model version, so a re-index only pays for chunks that actually changed.

Retrieval: dense alone is not enough

Dense retrieval is good at "documents about the same idea in different words." It is bad at exact tokens. Sparse retrieval — BM25 and friends — is the mirror image: it can't generalize past vocabulary, but if the user typed MX-7741-B, it will find the document containing MX-7741-B and rank it first.

Real queries contain both. "Why does the webhook retry loop throw ERR_CONN_RESET on the billing endpoint?" needs semantic matching on webhook retry loop and literal matching on the error code. So run both and fuse.

Reciprocal rank fusion is the right default because it fuses ranks, not scores. Cosine similarities and BM25 scores live on incompatible scales with corpus-dependent distributions; normalizing them into agreement is a tuning problem that never quite stays solved. RRF sidesteps it entirely and has one constant.

scores = defaultdict(float)
for ids, weight in zip(ranked_lists, weights):
    for rank, chunk_id in enumerate(ids):
        scores[chunk_id] += weight / (k + rank + 1)  # k damps the top ranks; 60 is the usual default

That is the whole algorithm. Each list contributes a score that decays with position, a chunk both legs ranked well accumulates from both, and the fused order is scores sorted descending. Around it: embed the query, run the dense and the sparse leg with the same filters and the same candidate limit, fuse the two ID lists, and hydrate the top of the fused list back into chunks — call that hybrid_search.

Two operational notes. First, both legs must apply the same metadata filters — an ACL enforced on one leg and not the other is a data leak, not a ranking bug. Second, the candidate limit should be generous. This stage is a recall stage; precision is the reranker's job.

Query transformation: useful, but not for free

Every transformation trades latency for recall, and each one is a separate model call before retrieval even starts.

  • Multi-query — generate a few paraphrases, retrieve for each, fuse with RRF. This is the one that earns its keep most consistently, because it patches vocabulary mismatch, and the retrievals parallelize so you pay one model call plus the slowest retrieval.
  • Query decomposition — split a compound question ("how does X compare to Y under Z?") into sub-questions retrieved independently. Essential when a single query genuinely cannot match any single chunk. Wasteful and occasionally harmful on simple lookups, because a decomposer handed an atomic question will invent sub-questions that drift.
  • HyDE — have the model hallucinate a plausible answer, embed that, and search with it. The idea is sound: a fake answer lives in the same region of embedding space as a real one, which fixes the shape mismatch between short questions and long passages. In practice it's the most expensive option and it degrades badly in specialized domains, where the hallucinated document is confidently wrong in exactly the vocabulary that matters.

Route rather than always-on. Short keyword-ish queries go straight to hybrid retrieval; long or compound ones earn a transformation. An always-on HyDE stage on a latency-sensitive product is a self-inflicted wound.

Reranking: paying for precision

A bi-encoder embeds the query and the document separately, so nothing in the document vector was computed with any knowledge of the query. A cross-encoder takes the query and passage together and scores the pair jointly, which is why it's dramatically better at judging relevance — and why it can't be precomputed.

That's the whole tradeoff. There is no index to amortize the cost into: reranking N candidates is N forward passes at query time. It adds meaningful latency and it scales linearly with the candidate count, so the candidate count is your budget dial.

def retrieve(query, filters, *, candidates: int = 50, final_k: int = 6):
    pool = hybrid_search(query, filters=filters, limit=candidates)         # stage 1 — recall, index-backed
    scores = cross_encoder.predict([(query, c.embed_text) for c in pool])  # stage 2 — one forward pass per pair
    ranked = sorted(zip(pool, scores), key=lambda cs: cs[1], reverse=True)
    return [c for c, s in ranked if s >= RELEVANCE_FLOOR][:final_k]

Two things make this pay off:

  • Tune candidates, not the model. Going wider improves the ceiling on what the reranker can find, and costs latency linearly. Somewhere there's a knee where extra candidates stop containing anything relevant. Find it on your data and fix it.
  • Use the absolute score, not just the order. Ranking always returns something. A relevance floor is what lets your system distinguish "here are six good passages" from "nothing in the corpus answers this," and that distinction is the difference between a useful assistant and a confident liar.

Context assembly: the part everyone skips

You now have six good chunks. How you lay them out still changes the answer.

Position matters. Models attend most reliably to the beginning and the end of a long context and are measurably weaker in the middle — the "lost in the middle" effect. So don't paste your ranked list in rank order, which buries your second-best evidence in the worst position. Interleave: strongest first, next-strongest last, and fill inward. A ranked [c0, c1, c2, c3] goes into the prompt as [c0, c2, c3, c1].

Deduplicate. Overlapping chunks, near-identical boilerplate, and the same passage reached by both retrieval legs all waste context and, worse, act as a vote — three copies of a stale paragraph will outweigh one copy of the current one. Dedupe on content_hash, keeping the highest-ranked occurrence.

Cite structurally. Give each chunk a stable, short label in the prompt and instruct the model to reference those labels. Post-hoc citation matching (find which source a sentence came from after generation) is guesswork; labels make attribution mechanical and let you render real links. Render each block as [S1] Billing API / Webhooks followed by the chunk text, and keep a registry — {"S1": chunk} — beside the assembled string.

The registry maps labels back to chunks, so when the model writes "the timeout is 30s [S3]", you can resolve S3 to a real URI and page number without parsing prose. Instruct the model explicitly to answer only from the labelled sources and to say so when they don't contain the answer — grounding is a prompt contract, not something retrieval gives you for free.

The vector store is an operational system

This is where RAG projects quietly rot. The index worked on the day you built it and nobody owns it after that.

Index type

HNSW IVF (+ PQ)
Structure Navigable small-world graph Cluster the space, search the nearest cells
Build Slower, incremental Faster, but needs a training pass over representative vectors
Memory Higher — the graph itself is large Lower, especially with product quantization
Recall/latency dial ef_search at query time nprobe at query time
Inserts Supported natively, graph degrades slowly Fine until the data distribution drifts from the trained centroids
Deletes Usually tombstones; needs periodic compaction Same

For most application-sized corpora, HNSW with sane M and ef_construction, tuned via ef_search, is the boring correct answer. Reach for IVF+PQ when vectors no longer fit comfortably in memory. And know which one you have, because the failure modes differ: HNSW degrades gradually as you churn the graph, while IVF degrades when your content changes character and the centroids no longer describe it — a corpus that grows into a new product line can lose recall on the new documents specifically.

Filtering is a correctness feature

Metadata filtering isn't a convenience; in any multi-tenant system it's the security boundary. The thing to understand is when the filter is applied.

  • Post-filtering retrieves top-k then discards non-matching results. Simple, and catastrophically wrong for selective filters: ask for one tenant's documents and you may get zero results back even though thousands exist, because none reached the global top-k.
  • Pre-filtering restricts the search to the matching subset. Correct, but naive implementations degrade into a scan when the subset is small.

Most mature stores do filtered graph traversal or per-tenant partitioning. Verify which yours does, because "returns nothing for small tenants" is a bug you will otherwise ship.

The filter every query carries is the same four keys:

  • tenant_id — the caller's tenant. A hard boundary, never optional.
  • acl_group — {"$in": user.groups}.
  • is_deleted: False — the soft-delete flag.
  • model_version — pinned to the active embedding model; never mix embedding spaces.

Partition by tenant at the collection or namespace level when the store supports it. It makes the boundary structural rather than a filter someone can forget to pass.

Keeping the index honest

The index is a cache of your source of truth, and like every cache it will drift.

  • Key by content, not by position. chunk_id as doc_id:ordinal plus a content_hash lets you diff a re-ingested document against what's indexed and touch only what changed.
  • Deletes must actually propagate. A document unpublished in the CMS but still in the index will be cited by your assistant, with a link, to a customer. This is the failure mode that turns into an incident. Soft-delete with a filtered flag for immediacy, hard-delete on compaction.
  • Re-ingest is delete-then-insert for the whole document. If a doc goes from 40 chunks to 37, upserting 37 leaves 3 orphans behind that still match queries.
  • Run a reconciliation job. Periodically walk the source of truth and the index and report divergence: documents in one and not the other, hash mismatches, wrong model version. Treat non-zero divergence as an alert, not a chore.
  • Rebuild behind an alias. Build the new index into a new collection, verify, then flip a pointer. Re-embedding in place with live traffic gives you an index that is half one model and half another, which fails quietly.

Measure retrieval separately from the answer

If you only score final answers, you learn that something is wrong and nothing about where. Retrieval needs its own numbers against a small labelled set of query-to-gold-chunk pairs — a few hundred, curated from real user queries, is enough to be useful.

The ones worth tracking: recall@k (is the gold chunk anywhere in the candidate pool?), which validates your recall stage and your chunking; MRR or nDCG (how high did it rank?), which validates fusion and reranking; and faithfulness (is every claim in the answer supported by a retrieved chunk?), which is the only one that catches a model contradicting good context. Track recall@k at the candidate stage and after reranking — a gap between them tells you the reranker is throwing away good evidence, which is a different fix from retrieving badly in the first place.

Debugging: two failures that look identical

A user reports a wrong answer. There are exactly two root causes and they have nothing in common:

  1. The right information never reached the model. Retrieval failure.
  2. The right information reached the model and it ignored, misread, or overrode it. Generation failure.

You cannot tell these apart from the answer. So don't try — separate them with one mechanical step: log the assembled context for every request, then read it. Was the correct passage in there or not? That single question splits the problem in half, and everything downstream follows from the answer.

If the passage wasn't there, walk backwards through the stages to find where it dropped out. Query the corpus directly for the gold chunk — does it even exist as a chunk? Then check whether hybrid retrieval surfaced it in the candidate pool, then whether the reranker kept it.

If it was there, the retrieval side is fine and you're debugging the prompt. The cleanest test is an ablation: hand the model hand-picked perfect context. If it still answers wrong, no amount of retrieval work will save you.

Symptom Likely cause Fix
Gold chunk isn't in the candidate pool at all Chunking split the fact, or no contextual prefix Structure-aware chunking; prepend heading path before embedding
Works for concepts, fails on IDs and error codes Dense-only retrieval Add BM25 leg and fuse with RRF
Candidate pool contains it, final context doesn't Reranker or final_k too aggressive Widen candidates, lower the relevance floor, compare recall@k before and after rerank
Zero results for some users, fine for others Post-filtering with a selective filter Move to pre-filtering or per-tenant namespaces
Answers cite documents that no longer exist Deletes never propagated Reconciliation job; delete-then-insert on re-ingest
Quality dropped overnight with no code change Mixed embedding spaces after a model swap Pin model_version in the filter; rebuild behind an alias
Context is correct, answer contradicts it Generation failure Ablate with hand-picked context; tighten the grounding instruction and citation contract
Model hedges or blends two answers Contradictory or duplicated chunks Deduplicate; prefer recency via metadata; surface the conflict instead of hiding it
Right facts, buried mid-context Lost in the middle Reorder so the strongest evidence sits at the edges; cut final_k

TL;DR checklist

  • Chunk along structure, not character counts; prepend the heading path before embedding
  • Store chunk text and metadata in a real database — the vector store is a rebuildable derived index
  • Stamp every vector with its embedding model version and filter on it
  • Always run hybrid: dense for meaning, BM25 for identifiers, fuse with RRF
  • Retrieve wide, rerank narrow; use the reranker's absolute score to allow an honest "I don't know"
  • Route query transformations by query shape instead of enabling them globally
  • Deduplicate, order for the edges of the context window, and cite with stable labels
  • Apply the same filters to every retrieval leg; make tenancy structural, not a parameter
  • Make deletes propagate and reconcile the index against the source of truth on a schedule
  • Log the assembled context on every request — it's the only way to split retrieval bugs from generation bugs

Parting thoughts

RAG isn't a model problem, and past the first week it's barely a prompting problem. It's an information retrieval system with a language model stapled to the end, and the discipline it rewards is the boring kind: knowing what's in your index, knowing when it last changed, and being able to answer "why did it return that?" for any query a user sends you.

Agentic setups that let a model plan its own multi-step searches build directly on top of this — they make the retrieval loop smarter, but they inherit every weakness in the pipeline underneath. Get the chunks, the fusion, and the index operations right first. Everything above them is leverage on that foundation.