October 4, 2026
Scaling Inference

RAG Architecture in 2027: Hybrid Retrieval, Rerankers, and Graph Indexes

RAG Architecture in 2027: Hybrid Retrieval, Rerankers, and Graph Indexes

A current RAG stack is a pipeline, not a vector search. Query understanding, metadata filtering, hybrid retrieval, reranking, and context assembly each remove a distinct failure mode. Naive top-k embedding search survives only in prototypes, because it fails on exact identifiers, fresh documents and multi-part questions.

Your retrieval is returning plausible but unhelpful chunks, and swapping the embedding model barely moved it. That is the situation almost every team reaches around month three, and it is a structural problem rather than a tuning one.

Single-stage vector search asks one component to do query interpretation, filtering, ranking and selection simultaneously. Splitting those responsibilities is what the last few years of practice actually amount to.

Figure 1 — The stages of a current retrieval pipeline

  • STAGE 1Query understandingRewrite, expand, decompose multi-part questions, resolve pronouns from history
  • STAGE 2Metadata filterNarrow by type, date, permission and tenant before any semantic work
  • STAGE 3Hybrid retrieveDense vectors plus keyword search, fused into one candidate set
  • STAGE 4RerankCross-encoder scores query and passage together, not separately
  • STAGE 5AssembleDeduplicate, order by score, place strongest nearest the question

Each stage removes a failure the previous one cannot see. Filtering cannot fix bad ranking, reranking cannot recover a document that was never retrieved, and assembly cannot compensate for a query the system misunderstood.

Why is hybrid retrieval close to mandatory?

Because dense and sparse retrieval fail on opposite inputs, and real queries contain both kinds.

Embeddings capture meaning and handle paraphrase well. They handle exact strings badly: a part number, an error code, a person’s surname or an internal acronym has no useful semantic neighbourhood.

Keyword search is the mirror image. It matches identifiers exactly and fails completely when the user’s vocabulary differs from the document’s.

QueryDense aloneKeyword aloneHybrid
“how do I stop the app crashing on startup”GoodWeakGood
“error ORA-01555”WeakExcellentExcellent
“what did Okafor decide about the Q3 rollout”PartialPartialGood
“cancellation policy”GoodGoodGood
“SKU 44718-B availability”PoorExcellentExcellent

Row two and row five are why hybrid retrieval is not optional in any system touching identifiers, and almost every enterprise corpus is full of them.

What does a reranker actually change?

It scores the query and the passage together rather than comparing two independently-computed vectors.

A bi-encoder embeds the query and the document separately, so it never sees them side by side. A cross-encoder reads both at once and can notice that a passage mentioning the right topic answers a different question about it.

The cost is that cross-encoders cannot pre-compute, so they only run over a candidate set. That is exactly why the pipeline shape is retrieve broadly, then rerank narrowly.

Watch for this

Reranking is frequently the single highest-return addition to an underperforming RAG system, and it is often skipped because teams reach for a better embedding model first. Retrieve 50 candidates, rerank to 10, and measure. If Recall@50 is high while Recall@10 is poor, the correct documents are already being found and simply ranked badly, which a reranker fixes directly.

When do graph indexes earn their cost?

Rarely, and specifically for corpus-wide questions no chunk can answer.

Microsoft’s GraphRAG work targets what it calls global sensemaking questions, reporting gains in comprehensiveness and diversity on datasets in the 1 million token range (arXiv:2404.16130). That is a genuine capability, and a narrow one.

Graph construction requires an LLM extraction call per text unit, which moves indexing from an embedding cost to an inference cost. For a system answering targeted questions, that expense buys nothing.

The honest guidance: add a graph layer when you have measured that a meaningful share of your traffic asks thematic questions about the corpus as a whole. Not before.

In what order should you add layers?

Cheapest and highest-yield first, measuring after each.

  1. Metadata filtering. Nearly free and frequently the largest single gain, because retrieving from the wrong document set cannot be fixed downstream.
  2. Hybrid retrieval. Solves the identifier failure completely.
  3. Reranking. The biggest quality gain per unit of effort in most systems.
  4. Query rewriting. Matters most for conversational interfaces where pronouns and ellipsis are common.
  5. Better chunking. Structural work that raises the ceiling for everything above it.
  6. Graph indexing. Only for corpus-wide sensemaking, and only once measured.

Notice what is absent from the top of that list: changing the embedding model. It belongs after all of the above, because the leaderboard leaders cluster tightly while these architectural gaps are large.

What breaks in production that never breaks in testing?

  • Multi-part questions. “What is the refund policy and how do I start one” needs decomposition into two retrievals, not one.
  • Conversational ellipsis. “What about for annual plans?” is unretrievable without rewriting against history.
  • Permission drift. A document’s access rules change and the index does not know, so filtering silently serves stale permissions.
  • Temporal ambiguity. “The latest policy” requires date-aware ranking that pure similarity cannot express.
  • Near-duplicate documents. Five versions of one policy fill the context window with the same content at slightly different vintages.

The last of these is worth explicit handling. Deduplication before assembly costs almost nothing and recovers window space that would otherwise be spent on repetition.

What do experienced teams do differently?

They log the whole pipeline, not just the final answer.

Storing the rewritten query, the filter applied, the candidate set, the rerank scores and the final assembled context turns “the answer was wrong” into a specific diagnosis: the filter excluded the right document, or the reranker demoted it, or it was never retrieved at all. Those need entirely different fixes and are indistinguishable from the output alone.

They also measure recall at each stage separately. Recall@50 before reranking and Recall@10 after tells you immediately whether you have a retrieval problem or a ranking problem, which is the single most useful diagnostic in the stack.

A short glossary

Bi-encoder
A model embedding query and document separately, enabling pre-computed vectors and fast search.
Cross-encoder
A model scoring query and passage together, more accurate and impossible to pre-compute.
Hybrid retrieval
Combining dense vector search with sparse keyword search so identifiers and paraphrase both work.
Query rewriting
Reformulating a user’s question, often against conversation history, before retrieval runs.
Rank fusion
Merging ranked lists from several retrievers into a single ordered candidate set.

Three things to do next

  1. Measure Recall@50 and Recall@10 separately on your existing eval set. The gap between them tells you whether to invest in retrieval or in ranking.
  2. Add keyword search alongside your vector search and fuse the results. This is a day of work and fixes every identifier query at once.
  3. Log the full pipeline for a week before changing anything else. Most teams discover their real failure is in a stage they were not looking at.

Key takeaways

  • Modern RAG is a multi-stage pipeline; single-stage vector search asks one component to do four jobs.
  • Dense and sparse retrieval fail on opposite inputs, which makes hybrid retrieval mandatory wherever identifiers appear.
  • Cross-encoder reranking is usually the highest-return addition to an underperforming system.
  • A large gap between Recall@50 and Recall@10 means you have a ranking problem, not a retrieval problem.
  • Graph indexes serve corpus-wide sensemaking questions and cost an LLM call per text unit to build.
  • Metadata filtering is nearly free and often the single largest gain available.
  • Changing the embedding model belongs near the end of the list, not the start.

Frequently asked questions

Is plain vector search still good enough for RAG?

Only for prototypes. It fails on exact identifiers, cannot express date or permission constraints, and asks a single component to handle query interpretation, filtering and ranking at once. Production systems separate those responsibilities into stages, each removing a failure the others cannot see.

What is the difference between a bi-encoder and a cross-encoder?

A bi-encoder embeds query and document separately, so vectors can be pre-computed and searched quickly. A cross-encoder reads both together and scores them jointly, which is more accurate but cannot be pre-computed. That is why cross-encoders rerank a candidate set rather than searching the corpus.

Do I need hybrid search if my embedding model is good?

Yes, if your corpus contains identifiers. Part numbers, error codes, SKUs and acronyms have no meaningful semantic neighbourhood, so no embedding model handles them reliably. Keyword search matches them exactly, and fusing both covers inputs neither handles alone.

How do I know whether to add a reranker?

Compare Recall@50 against Recall@10 on your evaluation set. If the first is high and the second is poor, the right documents are being retrieved and ranked badly, which is precisely what reranking fixes. If both are low, improve retrieval before ranking.

When is a graph index worth building?

When a measured share of your traffic asks questions about the corpus as a whole rather than about specific passages. Graph construction requires an LLM extraction call per text unit, so it is expensive, and it buys nothing for targeted factual lookup.

What should I fix first in an underperforming RAG system?

Metadata filtering, then hybrid retrieval, then reranking. All three are cheaper than changing your embedding model and typically deliver more. Retrieving from the wrong document set is a failure no downstream component can repair.

How do I handle follow-up questions in a conversation?

Rewrite the query against conversation history before retrieval. A question like “what about for annual plans” contains almost nothing retrievable on its own. Query rewriting resolves the pronouns and ellipsis into a self-contained question the retriever can actually match.

Why does my context window fill with near-duplicate documents?

Because similarity search happily returns several versions of the same document, all genuinely similar to the query. Deduplicate the candidate set before assembling context. It costs almost nothing and recovers window space otherwise spent on repeated content.

References

    Avatar photo
    Following her Bachelor's degree in Information Technology, Emma Hawkins actively participated in several student-led tech projects including the Cambridge Blockchain Society and graduated with top honors from the University of Cambridge. Emma, keen to learn more in the fast changing digital terrain, studied a postgraduate diploma in Digital Innovation at Imperial College London, focusing on sustainable tech solutions, digital transformation strategies, and newly emerging technologies.Emma, with more than ten years of technological expertise, offers a well-rounded skill set from working in many spheres of the company. Her path of work has seen her flourish in energetic startup environments, where she specialized in supporting creative ideas and hastening blockchain, Internet of Things (IoT), and smart city technologies product development. Emma has played a range of roles from tech analyst, where she conducted thorough market trend and emerging innovation research, to product manager—leading cross-functional teams to bring disruptive products to market.Emma currently offers careful analysis and thought leadership for a variety of clients including tech magazines, startups, and trade conferences using her broad background as a consultant and freelancing tech writer. Making creative technology relevant and understandable to a wide spectrum of listeners drives her in bridging the gap between technical complexity and daily influence. Emma is also highly sought for as a speaker at tech events where she provides her expertise on IoT integration, blockchain acceptance, and the critical role sustainability plays in tech innovation.Emma regularly attends conferences, meetings, and web forums, so becoming rather active in the tech community outside of her company. Especially interests her how technology might support sustainable development and environmental preservation. Emma enjoys trekking the scenic routes of the Lake District, snapping images of the natural beauties, and, in her personal time, visiting tech hotspots all around the world.

      Leave a Reply

      Your email address will not be published. Required fields are marked *