Sketio

AWS RAG Architecture with Amazon Bedrock

Updated

Retrieval-augmented generation (RAG) lets a language model answer from your own documents. Instead of retraining the model, the application looks up the passages that are relevant to the question and gives them to the model along with it. The answer can then point back to its sources.

Amazon Bedrock Knowledge Bases handles much of the pipeline. It reads your documents from a data source such as Amazon S3, splits them into chunks, turns the chunks into embeddings with an embedding model and stores them in a vector store. At question time it searches that store and can pass the best passages to a foundation model. This template shows both halves: loading documents, and answering a question.

sync and ingestembed / searchquestionS3 documentsKnowledge baseOpenSearch ServerlessUserApplicationFoundation modelRetrieveAndGeneratequestion + passages

Scroll sideways to see the whole diagram

AWS RAG Architecture with Amazon Bedrock. Open it in Sketio to change it.

Start from this diagram and edit it on your own board.

By continuing, you agree to the Terms of Service and Privacy Policy, including sending images of your strokes, diagram labels and similar data to providers in the United States (Cloudflare, Inc. and TypeSafe AI, Inc.) for AI conversion.

What each part does

S3 documents
The data source: a bucket that holds the files you want to answer from. When the files change, you sync the data source and the knowledge base ingests the changes.
Knowledge base
The Amazon Bedrock resource that ties the pieces together. During ingestion it splits documents into chunks, converts them to embeddings with the embedding model you chose and writes them to the vector store. At query time it embeds the question and searches the store.
OpenSearch Serverless
The vector store, one of several the knowledge base supports. It holds each chunk's embedding together with the chunk text and metadata, and returns the chunks closest in meaning to the question.
Foundation model
The model that writes the answer. It receives the question plus the retrieved passages as context. Pick one that Amazon Bedrock offers in your Region.
Application
Your code, for example an AWS Lambda function behind an API. It calls the knowledge base with the user's question and returns the result to the user.
User
The person asking a question in your chat or search interface.

How a question is answered

  1. Before any question: documents are uploaded to the S3 bucket, and a sync of the data source starts ingestion.
  2. The knowledge base splits the documents into chunks, converts each chunk to an embedding and stores it in the vector store, with a link back to the source document.
  3. A user asks a question in your application.
  4. The application calls the RetrieveAndGenerate API of Amazon Bedrock for the knowledge base, with the question and the model to use. The Retrieve API returns only the matching chunks, if you want to build the prompt yourself.
  5. The knowledge base converts the question to an embedding and searches the vector store for the chunks most similar to it.
  6. The question and those chunks go to the foundation model, which writes an answer.
  7. RetrieveAndGenerate returns the answer with citations that point to the source passages, and the application shows it to the user.

When to use it

Common variations

Choose another vector store

Besides OpenSearch Serverless, Amazon Bedrock Knowledge Bases supports Amazon S3 Vectors, Amazon OpenSearch Service managed clusters, Amazon Aurora PostgreSQL, Amazon Neptune Analytics, Pinecone, Redis Enterprise Cloud and MongoDB Atlas. Your choice of embedding model and vector size can narrow the options.

Change how documents are chunked

Chunking decides what a retrieved passage looks like. Try smaller chunks if answers pull in unrelated text, and larger ones if they lose context. The knowledge base offers several chunking strategies.

Filter by metadata

Attach metadata to your documents, such as department or year, and filter on it when you query, so people only get passages that apply to them.

Add guardrails

Amazon Bedrock Guardrails can check the question and the generated answer. They apply to the input and the model's response, not to the passages retrieved from the knowledge base, so keep sensitive content out of the data source.

Make it yours

Point the bucket at your own documents first, then choose the embedding model, the vector store and the answer model that fit your data, language and budget.

Opens this diagram as a board you can edit.

By continuing, you agree to the Terms of Service and Privacy Policy, including sending images of your strokes, diagram labels and similar data to providers in the United States (Cloudflare, Inc. and TypeSafe AI, Inc.) for AI conversion.

All templates