how-rag-works

How RAG Works

how-rag-works

RAG (Retrieval-Augmented Generation) connects an LLM to an external knowledge base so it can answer using information beyond what it was trained on. Here’s how the pieces in our graphic fit together:

Indexing phase (done ahead of time)

Documents → Chunked → Chunks: Source documents (PDFs, docs, web pages, etc.) get split into smaller chunks — typically a few hundred tokens each. Chunking matters because embedding a whole document loses granularity; you want retrieval to return focused, relevant passages, not entire files. Chunks → Embedded (via Embedding model) → Embedding → stored in Vector store: Each chunk is passed through an embedding model, which converts the text into a numeric vector that captures its semantic meaning. Those vectors (plus a reference to the original text) are stored in a vector store/database, indexed for fast similarity search.

Query phase (happens at runtime, per user query)

User query → Embedding model: When a user asks a question, that query is also run through the same embedding model, turning it into a vector in the same space as the stored chunks. Query embedding → Vector store → retrieves relevant chunks: The vector store compares the query vector against stored chunk vectors (using cosine similarity or similar) and returns the top-k most relevant chunks. Retrieved chunks → augment the prompt: the retrieved chunks are inserted into the prompt alongside the user’s original query. That combined prompt (query + retrieved context) is what gets sent to the LLM.

LLM generates a response: The LLM reads the augmented prompt and generates an answer grounded in the retrieved content, rather than relying solely on its training data.

Overall Flow:

how-rag-works-flow


The key insight: the embedding model is used twice — once to index the documents, once to embed the incoming query — and it’s the vector store’s retrieval result (not the embedding model itself) that augments what the LLM sees.