A current RAG stack is a pipeline, not a vector search. Query understanding, metadata filtering, hybrid retrieval, reranking, and context assembly each remove a distinct failure mode. Naive top-k embedding search survives only in prototypes, because it fails on exact identifiers, fresh documents and multi-part questions.
Your retrieval is returning plausible but unhelpful chunks, and swapping the embedding model barely moved it. That is the situation almost every team reaches around month three, and it is a structural problem rather than a tuning one.
Single-stage vector search asks one component to do query interpretation, filtering, ranking and selection simultaneously. Splitting those responsibilities is what the last few years of practice actually amount to.
Figure 1 — The stages of a current retrieval pipeline
- STAGE 1Query understandingRewrite, expand, decompose multi-part questions, resolve pronouns from history
- STAGE 2Metadata filterNarrow by type, date, permission and tenant before any semantic work
- STAGE 3Hybrid retrieveDense vectors plus keyword search, fused into one candidate set
- STAGE 4RerankCross-encoder scores query and passage together, not separately
- STAGE 5AssembleDeduplicate, order by score, place strongest nearest the question
Each stage removes a failure the previous one cannot see. Filtering cannot fix bad ranking, reranking cannot recover a document that was never retrieved, and assembly cannot compensate for a query the system misunderstood.
Why is hybrid retrieval close to mandatory?
Because dense and sparse retrieval fail on opposite inputs, and real queries contain both kinds.
Embeddings capture meaning and handle paraphrase well. They handle exact strings badly: a part number, an error code, a person’s surname or an internal acronym has no useful semantic neighbourhood.
Keyword search is the mirror image. It matches identifiers exactly and fails completely when the user’s vocabulary differs from the document’s.
| Query | Dense alone | Keyword alone | Hybrid |
|---|---|---|---|
| “how do I stop the app crashing on startup” | Good | Weak | Good |
| “error ORA-01555” | Weak | Excellent | Excellent |
| “what did Okafor decide about the Q3 rollout” | Partial | Partial | Good |
| “cancellation policy” | Good | Good | Good |
| “SKU 44718-B availability” | Poor | Excellent | Excellent |
Row two and row five are why hybrid retrieval is not optional in any system touching identifiers, and almost every enterprise corpus is full of them.
What does a reranker actually change?
It scores the query and the passage together rather than comparing two independently-computed vectors.
A bi-encoder embeds the query and the document separately, so it never sees them side by side. A cross-encoder reads both at once and can notice that a passage mentioning the right topic answers a different question about it.
The cost is that cross-encoders cannot pre-compute, so they only run over a candidate set. That is exactly why the pipeline shape is retrieve broadly, then rerank narrowly.
Watch for this
Reranking is frequently the single highest-return addition to an underperforming RAG system, and it is often skipped because teams reach for a better embedding model first. Retrieve 50 candidates, rerank to 10, and measure. If Recall@50 is high while Recall@10 is poor, the correct documents are already being found and simply ranked badly, which a reranker fixes directly.
When do graph indexes earn their cost?
Rarely, and specifically for corpus-wide questions no chunk can answer.
Microsoft’s GraphRAG work targets what it calls global sensemaking questions, reporting gains in comprehensiveness and diversity on datasets in the 1 million token range (arXiv:2404.16130). That is a genuine capability, and a narrow one.
Graph construction requires an LLM extraction call per text unit, which moves indexing from an embedding cost to an inference cost. For a system answering targeted questions, that expense buys nothing.
The honest guidance: add a graph layer when you have measured that a meaningful share of your traffic asks thematic questions about the corpus as a whole. Not before.
In what order should you add layers?
Cheapest and highest-yield first, measuring after each.
- Metadata filtering. Nearly free and frequently the largest single gain, because retrieving from the wrong document set cannot be fixed downstream.
- Hybrid retrieval. Solves the identifier failure completely.
- Reranking. The biggest quality gain per unit of effort in most systems.
- Query rewriting. Matters most for conversational interfaces where pronouns and ellipsis are common.
- Better chunking. Structural work that raises the ceiling for everything above it.
- Graph indexing. Only for corpus-wide sensemaking, and only once measured.
Notice what is absent from the top of that list: changing the embedding model. It belongs after all of the above, because the leaderboard leaders cluster tightly while these architectural gaps are large.
What breaks in production that never breaks in testing?
- Multi-part questions. “What is the refund policy and how do I start one” needs decomposition into two retrievals, not one.
- Conversational ellipsis. “What about for annual plans?” is unretrievable without rewriting against history.
- Permission drift. A document’s access rules change and the index does not know, so filtering silently serves stale permissions.
- Temporal ambiguity. “The latest policy” requires date-aware ranking that pure similarity cannot express.
- Near-duplicate documents. Five versions of one policy fill the context window with the same content at slightly different vintages.
The last of these is worth explicit handling. Deduplication before assembly costs almost nothing and recovers window space that would otherwise be spent on repetition.
What do experienced teams do differently?
They log the whole pipeline, not just the final answer.
Storing the rewritten query, the filter applied, the candidate set, the rerank scores and the final assembled context turns “the answer was wrong” into a specific diagnosis: the filter excluded the right document, or the reranker demoted it, or it was never retrieved at all. Those need entirely different fixes and are indistinguishable from the output alone.
They also measure recall at each stage separately. Recall@50 before reranking and Recall@10 after tells you immediately whether you have a retrieval problem or a ranking problem, which is the single most useful diagnostic in the stack.
A short glossary
- Bi-encoder
- A model embedding query and document separately, enabling pre-computed vectors and fast search.
- Cross-encoder
- A model scoring query and passage together, more accurate and impossible to pre-compute.
- Hybrid retrieval
- Combining dense vector search with sparse keyword search so identifiers and paraphrase both work.
- Query rewriting
- Reformulating a user’s question, often against conversation history, before retrieval runs.
- Rank fusion
- Merging ranked lists from several retrievers into a single ordered candidate set.
Three things to do next
- Measure Recall@50 and Recall@10 separately on your existing eval set. The gap between them tells you whether to invest in retrieval or in ranking.
- Add keyword search alongside your vector search and fuse the results. This is a day of work and fixes every identifier query at once.
- Log the full pipeline for a week before changing anything else. Most teams discover their real failure is in a stage they were not looking at.
Key takeaways
- Modern RAG is a multi-stage pipeline; single-stage vector search asks one component to do four jobs.
- Dense and sparse retrieval fail on opposite inputs, which makes hybrid retrieval mandatory wherever identifiers appear.
- Cross-encoder reranking is usually the highest-return addition to an underperforming system.
- A large gap between Recall@50 and Recall@10 means you have a ranking problem, not a retrieval problem.
- Graph indexes serve corpus-wide sensemaking questions and cost an LLM call per text unit to build.
- Metadata filtering is nearly free and often the single largest gain available.
- Changing the embedding model belongs near the end of the list, not the start.
Frequently asked questions
Is plain vector search still good enough for RAG?
Only for prototypes. It fails on exact identifiers, cannot express date or permission constraints, and asks a single component to handle query interpretation, filtering and ranking at once. Production systems separate those responsibilities into stages, each removing a failure the others cannot see.
What is the difference between a bi-encoder and a cross-encoder?
A bi-encoder embeds query and document separately, so vectors can be pre-computed and searched quickly. A cross-encoder reads both together and scores them jointly, which is more accurate but cannot be pre-computed. That is why cross-encoders rerank a candidate set rather than searching the corpus.
Do I need hybrid search if my embedding model is good?
Yes, if your corpus contains identifiers. Part numbers, error codes, SKUs and acronyms have no meaningful semantic neighbourhood, so no embedding model handles them reliably. Keyword search matches them exactly, and fusing both covers inputs neither handles alone.
How do I know whether to add a reranker?
Compare Recall@50 against Recall@10 on your evaluation set. If the first is high and the second is poor, the right documents are being retrieved and ranked badly, which is precisely what reranking fixes. If both are low, improve retrieval before ranking.
When is a graph index worth building?
When a measured share of your traffic asks questions about the corpus as a whole rather than about specific passages. Graph construction requires an LLM extraction call per text unit, so it is expensive, and it buys nothing for targeted factual lookup.
What should I fix first in an underperforming RAG system?
Metadata filtering, then hybrid retrieval, then reranking. All three are cheaper than changing your embedding model and typically deliver more. Retrieving from the wrong document set is a failure no downstream component can repair.
How do I handle follow-up questions in a conversation?
Rewrite the query against conversation history before retrieval. A question like “what about for annual plans” contains almost nothing retrievable on its own. Query rewriting resolves the pronouns and ellipsis into a self-contained question the retriever can actually match.
Why does my context window fill with near-duplicate documents?
Because similarity search happily returns several versions of the same document, all genuinely similar to the query. Deduplicate the candidate set before assembling context. It costs almost nothing and recovers window space otherwise spent on repeated content.
References
- Lewis, P., et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. arXiv:2005.11401.
- Edge, D., et al. From Local to Global: A Graph RAG Approach to Query-Focused Summarization. arXiv:2404.16130.
- Thakur, N., Reimers, N., Rücklé, A., Srivastava, A., and Gurevych, I. BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models. arXiv:2104.08663.
- Liu, N. F., et al. Lost in the Middle: How Language Models Use Long Contexts. arXiv:2307.03172. Basis for the ordering step in context assembly.
