When a retrieval-augmented generation system answers badly, what you see is a sentence, so the sentence gets blamed and the proposed fix is a better prompt or a larger model. The argument here is that most of these failures start earlier, in the corpus and the retriever, and that it is more useful to treat the whole thing as an information-retrieval system with a language model at the end. Below we split the pipeline into stages that fail in different ways, look at segmenting documents by their structure instead of by token count, go through where lexical and dense similarity each break, and show how to evaluate retrieval separately from answer quality.
RAG pipeline architecture: six stages and their failure modes
The pattern introduced by Lewis et al. (2020) separates parametric memory (what the model's weights encode) from non-parametric memory held in an external index, and conditions generation on documents fetched at inference time. That separation makes knowledge editable without retraining. It also gives the system six stages, each with its own failure mode.
The corpus is the source documents plus their governance: provenance, permissions, freshness, deletion. Segmentation cuts documents into retrievable units. Representation maps those units into index structures, which can be sparse lexical postings, dense vectors or both. Retrieval selects candidates. Reranking reorders them with a more expensive relevance model. Synthesis conditions the generator on the passages that survive, ideally with attribution.
The generator is a transformer (Vaswani et al., 2017) whose attention works over whatever tokens it is given. It has no way of wanting information it was not handed. If the correct passage is missing from the context window, no decoding strategy will bring it back. This is what makes diagnosis awkward: a defect at stage two shows up at stage six, where it is easy to mistake for a defect of the model.
Document chunking strategies and the structure they destroy
Chunking is usually treated as a tuning parameter, but it also flattens the document's structure into a list of strings, and the structure it throws away is often what made the document readable in the first place.
Fixed windows and tabular documents
A fixed window with chunk length c and stride s produces overlap o = c − s and cuts at token boundaries that have nothing to do with the document's own boundaries. Running prose tolerates this, because prose is locally redundant and the overlap repairs most severed sentences. Tables are a different matter. A table's meaning lives in the join between the header row and each body row, so a window that captures rows 40 through 60 of a price list without the header gives you a chunk in which the number 1,080 refers to nothing identifiable. The chunk is still perfectly retrievable, so the answer comes back confident, with a real source attached, and nothing looks missing.
Heading hierarchy and unresolved references
Specifications, contracts and manuals resolve references upward. A clause reading "this limit applies only to the configurations listed above" means nothing once it is separated from the heading and list that scoped it. Fixed-size chunking cuts those links at random, and anaphora such as "it", "the party" or "this process" can no longer be resolved.
The fix is to segment along the document's own structure. Split at headings, tables and clause boundaries, and prepend the heading path so the ancestor context travels with each piece. Keep tables whole, or serialise each row as a self-describing record that repeats the column names. When a small unit matches best but answers poorly, index the small unit and pass its enclosing parent to the generator. For a given corpus, the thing to find out is how small a span can get and still be true when read alone.
Lexical versus dense retrieval: where each notion of similarity fails
Lexical retrieval scores descend from the vector space model and the SMART tradition (Salton, 1971) and from probabilistic relevance weighting (Robertson & Sparck Jones, 1976). Their modern form is BM25: a term-frequency score, damped by saturation, weighted by inverse document frequency and normalised for length. BM25 treats relevance as term overlap, weighted towards the strings that are rare, and has no model of meaning at all.
Dense retrieval rests on the distributional hypothesis. Mikolov et al. (2013) made it practical for words, and contextual encoders later extended it to passages. Query and passage are each mapped to a vector and scored by cosine similarity, sim(q, d) = (q · d) / (‖q‖‖d‖). Here relevance means proximity in a learned space, which is a good but imperfect proxy for what a passage is about. Neither method matches what the user means by relevant, and they tend to fail on different queries.
| Query characteristic | Lexical (BM25) | Dense (embeddings) |
|---|---|---|
| Rare identifier (part code, SKU, error number) | Strong: rarity raises IDF weight | Weak: out-of-distribution token |
| Paraphrase with no shared terms | Weak: zero overlap scores zero | Strong: the trained-for case |
| Negation, e.g. "where X does not apply" | Weak: matches the affirmative | Weak: both forms sit close |
| Numerals and quantities | Partial: matches the literal string | Weak: magnitude poorly encoded |
| Cross-lingual query and document | Fails without translation | Feasible with a multilingual encoder |
| Low-frequency domain vocabulary | Strong | Degrades if the domain was unseen |
Two rows keep surprising teams. Rare identifiers are where dense retrieval is weakest and lexical retrieval strongest, and enterprise queries are full of them. Negation is where both fail. A sentence and its negation share nearly all their content words and sit at high cosine similarity, so the retriever cannot tell "the warranty covers water damage" from "the warranty does not cover water damage". Index tuning will not fix this, because the limitation is in the similarity measure itself. It has to be handled at reranking, or by declining to answer.
Hybrid retrieval and reranking
Since the two retrievers fail on different queries, running both and fusing their results beats either one alone on a mixed workload. Fusion can work on scores, which means normalising two incomparable scales, or on ranks, which does not. Rank-based reciprocal fusion is the safer default:
score(d) = Σ_i 1 / (k + rank_i(d))
where rank_i(d) is the position of document d in retriever i's list and k is a small constant (60 is a common choice) that limits how much the top of any single list can dominate. What matters is that a document ranked highly by either retriever survives, so a rare-identifier match does not get drowned out by semantically plausible passages.
Reranking deals with a different limitation. A bi-encoder embeds query and passage separately, which is what lets you precompute the index, but it also means the two never interact before scoring. A cross-encoder concatenates the pair and scores it jointly, so terms in the query can interact with terms in the passage. It is much more accurate and far too slow to run over a whole corpus. So the architecture is a funnel. Retrieve broadly with cheap methods, rerank a much smaller set with the expensive one, and pass only a few passages on to the generator.
Here is a rough cost model on assumed figures: each retriever returns 50 candidates in 20 ms, and a cross-encoder scores a pair in 8 ms with 32 pairs batched. Reranking 100 fused candidates then takes four batches, close to 60 ms in total, which is a small fraction of generation latency. The figures are hypothetical and only meant to show the proportions.
Grounding and attribution as product requirements
Attribution is often added last, as a presentation feature. It works better as a requirement set at the start, because it changes everything upstream. If the product has to show which span of which document supports each claim, then chunks need stable identifiers, offsets and document versions, synthesis has to answer only from the supplied passages, and the system has to be able to abstain.
Abstention is the requirement teams leave out most often, and it has the largest effect on trust. Top-k retrieval is unconditional: a retriever returns its best five passages whether or not any of them is relevant. Somewhere a threshold has to turn "these are the best candidates" into "these are good enough to answer from", and below that threshold the correct output is that the corpus does not contain the answer. Document-level attribution is not enough either. Pointing a reader at a forty-page manual hands the checking back to them, and saving them that work was the reason for building the system.
Retrieval evaluation metrics: measuring recall before prompt work
The most useful habit here is to measure the retriever on its own labelled set before touching any prompts, and the arithmetic shows why. If recall@k on a representative query set is 0.7, then for 30% of queries the evidence never reaches the context, and answer accuracy is capped at 0.7 whatever generator you use. No prompt change will raise that cap.
The metrics are the standard ones (Manning, Raghavan & Schütze, 2008): recall@k as the ceiling on answerability, precision@k as a proxy for how diluted the context is, mean reciprocal rank for where the first relevant item lands, and nDCG where graded relevance exists. One label they miss is context sufficiency, a binary judgement on whether the retrieved set actually allows a correct answer. A set can have perfect recall of relevant documents and still be insufficient, for example when the answer needs two figures from two passages and only one of those passages was retrieved. Answer-side evaluation, which has its own difficulties covered in evaluating language-model systems in production, should run only once the retrieval numbers are known. A faithfulness score over queries the retriever mostly fails tells you very little.
Corpus governance: freshness and erasure
A RAG index is derived data. Kleppmann (2017) argues that derived data should be handled as a materialised view of a source of truth, rather than as a second source of truth, and that puts a RAG index under the same obligations as any other derived table. It needs an explicit schema for what a chunk record contains, and a derivation step that can be re-run without duplicating or corrupting what it already produced. Both are covered at length in data contracts and idempotency. Two consequences follow, and both are architectural.
The first is freshness. Every answer is conditioned on whichever version of a document happened to be indexed. If the version is not recorded per chunk and shown in the citation, the system will confidently cite a superseded policy and nobody will be able to tell whether the model got it wrong or the index was stale.
The second is deletion. Under Article 17 of Regulation (EU) 2016/679 a data subject may obtain erasure of personal data, and Turkey's KVKK, Law No. 6698, imposes comparable obligations. Deleting the source satisfies neither if the derived artefacts survive, and a RAG pipeline produces a lot of them: chunk records, embedding vectors, sparse postings, approximate-nearest-neighbour structures that may not support deletion without a rebuild, reranker caches, evaluation fixtures, and logs that contain retrieved text. Erasure has to propagate through the whole derivation chain, so the chain has to be recorded, and teams make or skip that decision on day one.
Our most constrained example is RelationCRM, which is still in development. It scores more than 150 personal relationships across five dimensions and can import WhatsApp history for sentiment analysis. A personal corpus does not get much more sensitive than that, so we settled governance before the retrieval design. Storage stays on the device, following the privacy-by-design constraints of on-device AI. Contacts are anonymised to identifiers of the form Contact-001 before any model call, and the design is aligned to GDPR and KVKK under a zero-retention policy. These constraints have a cost. Anonymisation removes the proper nouns that lexical retrieval scores most sharply, which shifts the load onto dense retrieval over pseudonymised text. The product also covers ten languages. A shared embedding space lets a query in one language reach passages in another, while lexical scoring cannot cross that boundary at all, which is the same complementarity described in the table above.
Practical decision rules for RAG design
Segment along document structure and prepend the heading path. If a chunk is untrue when read alone, fix the segmentation instead of tuning around it. Keep tables intact or repeat the column names on every row. Fuse lexical and dense retrieval by rank, and if you go with a dense-only index, accept that you are giving up rare-identifier queries. Add reranking before deciding you have a latency problem. Make abstention a proper output with an explicit threshold. Before writing a prompt, build a labelled retrieval set of a few hundred real queries, and report recall@k next to every claim about answer quality. From day one, record the document version and derivation lineage on every chunk.
Limitations
This is an argument from mechanism and from our own project work. It is not an empirical study. We have not measured what share of RAG failures comes from retrieval versus generation, so the claim that retrieval dominates is a hypothesis that fits our experience and nothing stronger. The table gives directions that are well established qualitatively, without effect sizes, and the ordering can flip on particular corpora and encoders. The latency figures are a worked example under stated assumptions, not a benchmark. Everything here assumes a retrieve-then-read pipeline. Iterative and query-decomposition designs fail in materially different ways and are not covered. Index selection, quantisation and approximate-nearest-neighbour settings are left out too, and each of them loses some recall of its own. The RelationCRM section describes design decisions, not a controlled comparison.
References
- Kleppmann, M. (2017). Designing data-intensive applications.
- Lewis, P., et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks.
- Manning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to information retrieval.
- Mikolov, T., et al. (2013). Distributed representations of words and phrases and their compositionality.
- Regulation (EU) 2016/679 (General Data Protection Regulation), Article 17.
- Robertson, S. E., & Sparck Jones, K. (1976). Relevance weighting of search terms.
- Salton, G. (1971). The SMART retrieval system: experiments in automatic document processing.
- Turkey, Law No. 6698 on the Protection of Personal Data (KVKK).
- Vaswani, A., et al. (2017). Attention is all you need.