AWS SQS Lambda Architecture with Dead-Letter Queue
Amazon SQS and AWS Lambda are a standard way to take slow or bursty work off the request path. A producer sends a message to a queue and moves on. Lambda reads messages from the queue in batches and processes them at its own pace.
The dead-letter queue (DLQ) is what makes this safe to run. A message that keeps failing is moved aside after a set number of attempts, instead of blocking the queue or being retried forever, and an alarm tells you it happened.
Scroll sideways to see the whole diagram
Start from this diagram and edit it on your own board.
By continuing, you agree to the Terms of Service and Privacy Policy, including sending images of your strokes, diagram labels and similar data to providers in the United States (Cloudflare, Inc. and TypeSafe AI, Inc.) for AI conversion.
What each part does
- Producer app
- Anything that creates work: a web backend, another Lambda function or an AWS service. It calls SendMessage and does not wait for the result.
- Source queue
- A standard or FIFO SQS queue that holds messages durably until they are processed. Its redrive policy names the dead-letter queue and the maxReceiveCount.
- Lambda consumer
- A function wired to the queue by an event source mapping. Lambda polls the queue, invokes the function with a batch of messages, and deletes them from the queue only when the function succeeds.
- DynamoDB
- Where the consumer writes its result. Any downstream store works. Make the write idempotent, because a standard queue delivers at least once.
- Dead-letter queue
- A second SQS queue that receives a message once it has been received maxReceiveCount times without being deleted. Keep failed messages long enough to investigate, and redrive them after the bug is fixed.
- CloudWatch alarm
- Watches the DLQ's ApproximateNumberOfMessagesVisible metric and notifies you, for example through SNS, as soon as it is above zero.
How a message flows
- The producer sends a message to the source queue and carries on without waiting.
- Lambda's event source mapping polls the queue and invokes the function with a batch of messages.
- If the function returns successfully, Lambda deletes the batch from the queue and the function's result is written to DynamoDB.
- If the function throws, the messages become visible again after the visibility timeout and are retried.
- Once a message has been received maxReceiveCount times, SQS moves it to the dead-letter queue.
- A CloudWatch alarm on the number of messages in the DLQ notifies you, so someone can inspect and redrive them.
When to use it
- Order or image processing that should not slow down the web request that triggered it.
- Smoothing out spikes: the queue absorbs a burst and Lambda drains it at a limited concurrency.
- Calling a flaky downstream API, where failed calls should be retried a few times and then set aside.
Common variations
Use a FIFO queue
Choose FIFO when order matters or duplicates must be avoided. Messages in the same message group are processed in order. A FIFO queue needs a FIFO dead-letter queue.
Report partial batch failures
Return the IDs of the messages that failed (ReportBatchItemFailures) so the ones that succeeded are not retried together with them.
Fan out with SNS first
Publish to an SNS topic and subscribe several queues to it, so each independent consumer gets its own copy and its own dead-letter queue.
Tune the timings
Set the queue's visibility timeout to at least six times the function timeout, and give the DLQ a longer retention period than the source queue, because the clock starts when a message is first sent.
Make it yours
Rename the producer and the consumer after your own services, and replace DynamoDB with wherever the consumer writes its result.
Opens this diagram as a board you can edit.
By continuing, you agree to the Terms of Service and Privacy Policy, including sending images of your strokes, diagram labels and similar data to providers in the United States (Cloudflare, Inc. and TypeSafe AI, Inc.) for AI conversion.