AWS RAG Architecture with Amazon Bedrock
Retrieval-augmented generation (RAG) lets a language model answer from your own documents. Instead of retraining the model, the application looks up the passages that are relevant to the question and gives them to the model along with it. The answer can then point back to its sources.
Amazon Bedrock Knowledge Bases handles much of the pipeline. It reads your documents from a data source such as Amazon S3, splits them into chunks, turns the chunks into embeddings with an embedding model and stores them in a vector store. At question time it searches that store and can pass the best passages to a foundation model. This template shows both halves: loading documents, and answering a question.
Scroll sideways to see the whole diagram
Start from this diagram and edit it on your own board.
By continuing, you agree to the Terms of Service and Privacy Policy, including sending images of your strokes, diagram labels and similar data to providers in the United States (Cloudflare, Inc. and TypeSafe AI, Inc.) for AI conversion.
What each part does
- S3 documents
- The data source: a bucket that holds the files you want to answer from. When the files change, you sync the data source and the knowledge base ingests the changes.
- Knowledge base
- The Amazon Bedrock resource that ties the pieces together. During ingestion it splits documents into chunks, converts them to embeddings with the embedding model you chose and writes them to the vector store. At query time it embeds the question and searches the store.
- OpenSearch Serverless
- The vector store, one of several the knowledge base supports. It holds each chunk's embedding together with the chunk text and metadata, and returns the chunks closest in meaning to the question.
- Foundation model
- The model that writes the answer. It receives the question plus the retrieved passages as context. Pick one that Amazon Bedrock offers in your Region.
- Application
- Your code, for example an AWS Lambda function behind an API. It calls the knowledge base with the user's question and returns the result to the user.
- User
- The person asking a question in your chat or search interface.
How a question is answered
- Before any question: documents are uploaded to the S3 bucket, and a sync of the data source starts ingestion.
- The knowledge base splits the documents into chunks, converts each chunk to an embedding and stores it in the vector store, with a link back to the source document.
- A user asks a question in your application.
- The application calls the RetrieveAndGenerate API of Amazon Bedrock for the knowledge base, with the question and the model to use. The Retrieve API returns only the matching chunks, if you want to build the prompt yourself.
- The knowledge base converts the question to an embedding and searches the vector store for the chunks most similar to it.
- The question and those chunks go to the foundation model, which writes an answer.
- RetrieveAndGenerate returns the answer with citations that point to the source passages, and the application shows it to the user.
When to use it
- A question-answering assistant over internal documentation, policies or product manuals.
- A support tool that answers from a help centre and shows which article it used.
- Search that returns passages and an explanation instead of a list of links.
Common variations
Choose another vector store
Besides OpenSearch Serverless, Amazon Bedrock Knowledge Bases supports Amazon S3 Vectors, Amazon OpenSearch Service managed clusters, Amazon Aurora PostgreSQL, Amazon Neptune Analytics, Pinecone, Redis Enterprise Cloud and MongoDB Atlas. Your choice of embedding model and vector size can narrow the options.
Change how documents are chunked
Chunking decides what a retrieved passage looks like. Try smaller chunks if answers pull in unrelated text, and larger ones if they lose context. The knowledge base offers several chunking strategies.
Filter by metadata
Attach metadata to your documents, such as department or year, and filter on it when you query, so people only get passages that apply to them.
Add guardrails
Amazon Bedrock Guardrails can check the question and the generated answer. They apply to the input and the model's response, not to the passages retrieved from the knowledge base, so keep sensitive content out of the data source.
Make it yours
Point the bucket at your own documents first, then choose the embedding model, the vector store and the answer model that fit your data, language and budget.
Opens this diagram as a board you can edit.
By continuing, you agree to the Terms of Service and Privacy Policy, including sending images of your strokes, diagram labels and similar data to providers in the United States (Cloudflare, Inc. and TypeSafe AI, Inc.) for AI conversion.