Serverless Architecture Patterns
A comprehensive guide to designing, implementing, and optimising serverless applications using proven architectural patterns.
Overview
Serverless architecture enables developers to build and run applications without managing infrastructure, focusing on business logic whilst the cloud provider handles scaling, availability, and maintenance.
graph TB
subgraph "Serverless Architecture Overview"
Client[Client Application]
subgraph "API Layer"
APIGW[API Gateway]
end
subgraph "Compute Layer"
F1[Function A]
F2[Function B]
F3[Function C]
end
subgraph "Event Sources"
Q[Message Queue]
S[Object Storage]
DB[Database Stream]
SCH[Scheduler]
end
subgraph "Data Layer"
DDB[NoSQL Database]
RDS[Relational DB]
Cache[Cache Layer]
end
subgraph "Integration"
SNS[Notification Service]
SQS[Queue Service]
EventBus[Event Bus]
end
end
Client --> APIGW
APIGW --> F1
APIGW --> F2
Q --> F2
S --> F3
DB --> F1
SCH --> F3
F1 --> DDB
F2 --> RDS
F3 --> Cache
F1 --> SNS
F2 --> SQS
F3 --> EventBus
Function Composition
Key Concepts
Orchestration uses a central coordinator to manage workflow execution, whilst Choreography relies on decentralised, event-driven communication between services.
| Aspect | Orchestration | Choreography |
|---|---|---|
| Control | Centralised workflow engine | Decentralised event routing |
| Coupling | Tighter coupling to orchestrator | Loose coupling between services |
| Visibility | Clear workflow visualisation | Distributed tracing required |
| Error Handling | Centralised retry logic | Per-service error handling |
| Use Case | Complex workflows, transactions | Simple event reactions |
Orchestration Pattern
graph LR
subgraph "Step Functions Orchestration"
SF[Step Functions]
SF --> |"1. Validate"| F1[Validate Order]
F1 --> |"2. Process"| F2[Process Payment]
F2 --> |"3. Fulfil"| F3[Fulfil Order]
F3 --> |"4. Notify"| F4[Send Notification]
end
F2 --> |"Failure"| Comp[Compensation Logic]
Comp --> Refund[Refund Payment]
Common Patterns
AWS Step Functions (Orchestration)
{
"Comment": "Order processing workflow",
"StartAt": "ValidateOrder",
"States": {
"ValidateOrder": {
"Type": "Task",
"Resource": "arn:aws:lambda:region:account:function:validate",
"Next": "ProcessPayment",
"Catch": [{
"ErrorEquals": ["ValidationError"],
"Next": "OrderFailed"
}]
},
"ProcessPayment": {
"Type": "Task",
"Resource": "arn:aws:lambda:region:account:function:payment",
"Next": "FulfilOrder",
"Retry": [{
"ErrorEquals": ["PaymentRetryable"],
"MaxAttempts": 3,
"IntervalSeconds": 2,
"BackoffRate": 2
}]
},
"FulfilOrder": {
"Type": "Parallel",
"Branches": [
{
"StartAt": "UpdateInventory",
"States": {
"UpdateInventory": {
"Type": "Task",
"Resource": "arn:aws:lambda:region:account:function:inventory",
"End": true
}
}
},
{
"StartAt": "SendConfirmation",
"States": {
"SendConfirmation": {
"Type": "Task",
"Resource": "arn:aws:lambda:region:account:function:notify",
"End": true
}
}
}
],
"Next": "OrderComplete"
},
"OrderComplete": {
"Type": "Succeed"
},
"OrderFailed": {
"Type": "Fail",
"Error": "OrderProcessingFailed"
}
}
}
Choreography Pattern
graph LR
subgraph "Event-Driven Choreography"
O[Order Service] --> |"OrderCreated"| EB[Event Bus]
EB --> P[Payment Service]
P --> |"PaymentProcessed"| EB
EB --> I[Inventory Service]
I --> |"InventoryUpdated"| EB
EB --> N[Notification Service]
end
// Event publisher (Order Service)
const { EventBridgeClient, PutEventsCommand } = require('@aws-sdk/client-eventbridge');
const client = new EventBridgeClient({ region: 'eu-west-1' });
exports.handler = async (event) => {
const order = JSON.parse(event.body);
// Save order to database
await saveOrder(order);
// Publish event for choreography
const params = {
Entries: [{
Source: 'order.service',
DetailType: 'OrderCreated',
Detail: JSON.stringify({
orderId: order.id,
customerId: order.customerId,
items: order.items,
total: order.total,
timestamp: new Date().toISOString()
}),
EventBusName: 'orders-event-bus'
}]
};
await client.send(new PutEventsCommand(params));
return {
statusCode: 202,
body: JSON.stringify({ orderId: order.id, status: 'processing' })
};
};
Examples
Saga Pattern for Distributed Transactions
// Orchestrated saga with compensation
const sagaDefinition = {
steps: [
{
name: 'reserveInventory',
action: 'inventory:reserve',
compensation: 'inventory:release'
},
{
name: 'processPayment',
action: 'payment:charge',
compensation: 'payment:refund'
},
{
name: 'createShipment',
action: 'shipping:create',
compensation: 'shipping:cancel'
}
]
};
// Step Functions implementation with compensating transactions
const sagaStateMachine = {
StartAt: 'ReserveInventory',
States: {
ReserveInventory: {
Type: 'Task',
Resource: '${ReserveInventoryFunction}',
ResultPath: '$.inventoryReservation',
Next: 'ProcessPayment',
Catch: [{
ErrorEquals: ['States.ALL'],
ResultPath: '$.error',
Next: 'SagaFailed'
}]
},
ProcessPayment: {
Type: 'Task',
Resource: '${ProcessPaymentFunction}',
ResultPath: '$.paymentResult',
Next: 'CreateShipment',
Catch: [{
ErrorEquals: ['States.ALL'],
ResultPath: '$.error',
Next: 'ReleaseInventory'
}]
},
// ... compensation states
ReleaseInventory: {
Type: 'Task',
Resource: '${ReleaseInventoryFunction}',
Next: 'SagaFailed'
},
SagaFailed: {
Type: 'Fail',
Error: 'SagaExecutionFailed'
}
}
};
Event-Driven Design
Key Concepts
Event-driven architecture enables loose coupling between services through asynchronous message passing, improving scalability and resilience.
Core Principles:
- Event Sourcing: Store state changes as immutable events
- CQRS: Separate read and write models
- Eventual Consistency: Accept temporary inconsistencies for scalability
- Idempotency: Design handlers to safely process duplicate events
graph TB
subgraph "Event-Driven Architecture"
Producer[Event Producer]
subgraph "Event Router"
EB[Event Bus/Broker]
end
subgraph "Consumers"
C1[Consumer A]
C2[Consumer B]
C3[Consumer C]
end
subgraph "Event Store"
ES[(Event Log)]
end
end
Producer --> |"Publish Event"| EB
EB --> |"Route"| C1
EB --> |"Route"| C2
EB --> |"Route"| C3
Producer --> |"Persist"| ES
Common Patterns
Event Schema Definition
// CloudEvents specification compliant event
const orderEvent = {
specversion: '1.0',
type: 'com.company.order.created',
source: '/orders/service',
id: 'A234-1234-1234',
time: '2024-01-15T10:30:00Z',
datacontenttype: 'application/json',
data: {
orderId: 'ORD-12345',
customerId: 'CUST-67890',
items: [
{ productId: 'PROD-001', quantity: 2, price: 29.99 }
],
total: 59.98,
currency: 'GBP'
}
};
EventBridge Rule Configuration
# AWS SAM template for event routing
Resources:
OrderCreatedRule:
Type: AWS::Events::Rule
Properties:
Name: order-created-rule
EventBusName: !Ref OrderEventBus
EventPattern:
source:
- 'order.service'
detail-type:
- 'OrderCreated'
detail:
total:
- numeric: ['>=', 100]
Targets:
- Id: 'HighValueOrderProcessor'
Arn: !GetAtt HighValueOrderFunction.Arn
- Id: 'AnalyticsQueue'
Arn: !GetAtt AnalyticsQueue.Arn
# Dead Letter Queue for failed events
EventDLQ:
Type: AWS::SQS::Queue
Properties:
QueueName: event-dlq
MessageRetentionPeriod: 1209600 # 14 days
Idempotent Event Handler
const { DynamoDBClient } = require('@aws-sdk/client-dynamodb');
const { DynamoDBDocumentClient, GetCommand, PutCommand } = require('@aws-sdk/lib-dynamodb');
const client = DynamoDBDocumentClient.from(new DynamoDBClient({}));
exports.handler = async (event) => {
for (const record of event.Records) {
const eventData = JSON.parse(record.body);
const eventId = eventData.id;
// Check for duplicate processing (idempotency)
const existing = await client.send(new GetCommand({
TableName: 'ProcessedEvents',
Key: { eventId }
}));
if (existing.Item) {
console.log(`Event ${eventId} already processed, skipping`);
continue;
}
try {
// Process the event
await processEvent(eventData);
// Mark as processed with TTL
await client.send(new PutCommand({
TableName: 'ProcessedEvents',
Item: {
eventId,
processedAt: Date.now(),
ttl: Math.floor(Date.now() / 1000) + 86400 * 7 // 7 days
},
ConditionExpression: 'attribute_not_exists(eventId)'
}));
} catch (error) {
if (error.name === 'ConditionalCheckFailedException') {
console.log(`Concurrent processing detected for ${eventId}`);
continue;
}
throw error;
}
}
};
Examples
Fan-Out Pattern with SNS and SQS
graph LR
subgraph "Fan-Out Pattern"
Pub[Publisher] --> SNS[SNS Topic]
SNS --> Q1[SQS Queue 1]
SNS --> Q2[SQS Queue 2]
SNS --> Q3[SQS Queue 3]
Q1 --> F1[Function A]
Q2 --> F2[Function B]
Q3 --> F3[Function C]
end
# CloudFormation for fan-out pattern
Resources:
OrderTopic:
Type: AWS::SNS::Topic
Properties:
TopicName: order-events
InventoryQueue:
Type: AWS::SQS::Queue
Properties:
QueueName: inventory-updates
RedrivePolicy:
deadLetterTargetArn: !GetAtt InventoryDLQ.Arn
maxReceiveCount: 3
InventorySubscription:
Type: AWS::SNS::Subscription
Properties:
TopicArn: !Ref OrderTopic
Protocol: sqs
Endpoint: !GetAtt InventoryQueue.Arn
FilterPolicy:
eventType:
- OrderCreated
- OrderCancelled
Cold Start Mitigation
Key Concepts
Cold starts occur when a function instance must be initialised before handling a request, adding latency. Mitigation strategies include provisioned concurrency, optimised packaging, and architectural decisions.
Cold Start Components:
- Container Initialisation: Runtime environment setup
- Code Download: Fetching deployment package
- Runtime Initialisation: Language runtime startup
- Handler Initialisation: Application code initialisation
graph LR
subgraph "Cold Start Timeline"
A[Request] --> B[Container Init]
B --> C[Code Download]
C --> D[Runtime Init]
D --> E[Handler Init]
E --> F[Execution]
end
style B fill:#ff9999
style C fill:#ff9999
style D fill:#ffcc99
style E fill:#ffcc99
style F fill:#99ff99
Common Patterns
Provisioned Concurrency Configuration
# AWS SAM template
Resources:
CriticalFunction:
Type: AWS::Serverless::Function
Properties:
FunctionName: critical-api-handler
Handler: index.handler
Runtime: nodejs18.x
MemorySize: 1024
AutoPublishAlias: live
ProvisionedConcurrencyConfig:
ProvisionedConcurrentExecutions: 10
# Scheduled scaling for provisioned concurrency
ScalingTarget:
Type: AWS::ApplicationAutoScaling::ScalableTarget
Properties:
MaxCapacity: 50
MinCapacity: 5
ResourceId: !Sub function:${CriticalFunction}:live
ScalableDimension: lambda:function:ProvisionedConcurrency
ServiceNamespace: lambda
ScalingPolicy:
Type: AWS::ApplicationAutoScaling::ScalingPolicy
Properties:
PolicyName: utilisation-tracking
PolicyType: TargetTrackingScaling
ScalingTargetId: !Ref ScalingTarget
TargetTrackingScalingPolicyConfiguration:
TargetValue: 0.7
PredefinedMetricSpecification:
PredefinedMetricType: LambdaProvisionedConcurrencyUtilization
Optimised Function Packaging
// webpack.config.js for optimal bundling
const path = require('path');
const TerserPlugin = require('terser-webpack-plugin');
module.exports = {
mode: 'production',
target: 'node',
entry: './src/handler.js',
output: {
path: path.resolve(__dirname, 'dist'),
filename: 'index.js',
libraryTarget: 'commonjs2'
},
externals: [
// Keep AWS SDK external (available in Lambda runtime)
/^@aws-sdk\/.*/
],
optimization: {
minimise: true,
minimizer: [new TerserPlugin({
terserOptions: {
keep_fnames: true // Preserve function names for debugging
}
})]
},
resolve: {
extensions: ['.js', '.json']
}
};
Initialisation Optimisation
// Initialise outside handler (runs once per container)
const { DynamoDBClient } = require('@aws-sdk/client-dynamodb');
const { DynamoDBDocumentClient } = require('@aws-sdk/lib-dynamodb');
// Connection reuse - initialised once
const client = DynamoDBDocumentClient.from(
new DynamoDBClient({
maxAttempts: 3,
requestHandler: {
connectionTimeout: 3000,
socketTimeout: 3000
}
}),
{
marshallOptions: { removeUndefinedValues: true }
}
);
// Pre-compute static values
const CONFIG = {
tableName: process.env.TABLE_NAME,
region: process.env.AWS_REGION
};
// Handler function
exports.handler = async (event) => {
// Minimal work here - connections already established
return await processRequest(event, client, CONFIG);
};
Examples
Lambda SnapStart (Java)
// Enable SnapStart for Java functions
// AWS SAM template
/*
Resources:
JavaFunction:
Type: AWS::Serverless::Function
Properties:
Runtime: java17
SnapStart:
ApplyOn: PublishedVersions
AutoPublishAlias: live
*/
// Implement CRaC hooks for proper restoration
import org.crac.Context;
import org.crac.Core;
import org.crac.Resource;
public class Handler implements RequestHandler<APIGatewayProxyRequestEvent, APIGatewayProxyResponseEvent>, Resource {
private DynamoDbClient dynamoDbClient;
public Handler() {
// Register for checkpoint/restore notifications
Core.getGlobalContext().register(this);
// Initialise connections
this.dynamoDbClient = DynamoDbClient.create();
}
@Override
public void beforeCheckpoint(Context<? extends Resource> context) {
// Clean up resources before snapshot
// Close connections, clear caches with sensitive data
}
@Override
public void afterRestore(Context<? extends Resource> context) {
// Re-establish connections after restore
this.dynamoDbClient = DynamoDbClient.create();
}
@Override
public APIGatewayProxyResponseEvent handleRequest(
APIGatewayProxyRequestEvent event,
com.amazonaws.services.lambda.runtime.Context context) {
// Handle request
return new APIGatewayProxyResponseEvent()
.withStatusCode(200)
.withBody("Success");
}
}
Warming Strategy
// Scheduled warming function
const { LambdaClient, InvokeCommand } = require('@aws-sdk/client-lambda');
const lambda = new LambdaClient({});
const FUNCTIONS_TO_WARM = [
'critical-api-handler',
'payment-processor',
'order-validator'
];
exports.handler = async () => {
const warmingPromises = FUNCTIONS_TO_WARM.map(async (functionName) => {
const params = {
FunctionName: functionName,
InvocationType: 'RequestResponse',
Payload: JSON.stringify({ warmup: true })
};
try {
await lambda.send(new InvokeCommand(params));
console.log(`Warmed: ${functionName}`);
} catch (error) {
console.error(`Failed to warm ${functionName}:`, error.message);
}
});
await Promise.all(warmingPromises);
};
// In target function - handle warming requests
exports.targetHandler = async (event) => {
if (event.warmup) {
return { statusCode: 200, body: 'Warmed' };
}
// Normal processing
return await processRequest(event);
};
State Management Strategies
Key Concepts
Serverless functions are stateless by design. State must be externalised to databases, caches, or workflow services.
State Management Options:
| Service | Use Case | Latency | Cost Model |
|---|---|---|---|
| DynamoDB | Persistent state, high throughput | 1-10ms | Per request + storage |
| ElastiCache/Redis | Session data, caching | <1ms | Per hour |
| Step Functions | Workflow state | N/A | Per transition |
| S3 | Large objects, checkpoints | 50-100ms | Per request + storage |
| Parameter Store | Configuration | 10-50ms | Free tier available |
Common Patterns
Distributed Session Management
const Redis = require('ioredis');
// Connection outside handler for reuse
const redis = new Redis({
host: process.env.REDIS_HOST,
port: 6379,
connectTimeout: 5000,
maxRetriesPerRequest: 3
});
exports.handler = async (event) => {
const sessionId = event.headers['x-session-id'];
// Get session from Redis
let session = await redis.get(`session:${sessionId}`);
if (!session) {
// Create new session
session = {
id: sessionId,
created: Date.now(),
data: {}
};
} else {
session = JSON.parse(session);
}
// Update session
session.lastAccessed = Date.now();
session.data.requestCount = (session.data.requestCount || 0) + 1;
// Save with expiry (30 minutes)
await redis.setex(
`session:${sessionId}`,
1800,
JSON.stringify(session)
);
return {
statusCode: 200,
body: JSON.stringify({ session })
};
};
DynamoDB State Pattern
const { DynamoDBDocumentClient, UpdateCommand, GetCommand } = require('@aws-sdk/lib-dynamodb');
const { DynamoDBClient } = require('@aws-sdk/client-dynamodb');
const client = DynamoDBDocumentClient.from(new DynamoDBClient({}));
// Optimistic locking for state updates
exports.updateState = async (entityId, updateFn) => {
const maxRetries = 3;
for (let attempt = 0; attempt < maxRetries; attempt++) {
// Get current state
const current = await client.send(new GetCommand({
TableName: 'EntityState',
Key: { entityId }
}));
const currentVersion = current.Item?.version || 0;
const currentState = current.Item?.state || {};
// Apply update function
const newState = updateFn(currentState);
try {
// Conditional update with version check
await client.send(new UpdateCommand({
TableName: 'EntityState',
Key: { entityId },
UpdateExpression: 'SET #state = :state, #version = :newVersion, #updated = :updated',
ConditionExpression: 'attribute_not_exists(version) OR version = :currentVersion',
ExpressionAttributeNames: {
'#state': 'state',
'#version': 'version',
'#updated': 'updatedAt'
},
ExpressionAttributeValues: {
':state': newState,
':newVersion': currentVersion + 1,
':currentVersion': currentVersion,
':updated': new Date().toISOString()
}
}));
return newState;
} catch (error) {
if (error.name === 'ConditionalCheckFailedException') {
console.log(`Concurrent modification detected, retry ${attempt + 1}`);
continue;
}
throw error;
}
}
throw new Error('Max retries exceeded for state update');
};
Examples
Workflow State with Step Functions
graph TB
subgraph "Step Functions State Management"
Start([Start]) --> Init[Initialise State]
Init --> Process[Process Items]
Process --> Check{More Items?}
Check --> |Yes| Process
Check --> |No| Aggregate[Aggregate Results]
Aggregate --> End([End])
end
{
"Comment": "Stateful batch processing workflow",
"StartAt": "InitialiseState",
"States": {
"InitialiseState": {
"Type": "Pass",
"Result": {
"processedCount": 0,
"errors": [],
"results": []
},
"ResultPath": "$.state",
"Next": "ProcessBatch"
},
"ProcessBatch": {
"Type": "Map",
"ItemsPath": "$.items",
"MaxConcurrency": 10,
"ResultPath": "$.batchResults",
"Iterator": {
"StartAt": "ProcessItem",
"States": {
"ProcessItem": {
"Type": "Task",
"Resource": "${ProcessItemFunction}",
"End": true
}
}
},
"Next": "UpdateState"
},
"UpdateState": {
"Type": "Task",
"Resource": "arn:aws:states:::lambda:invoke",
"Parameters": {
"FunctionName": "${AggregateResultsFunction}",
"Payload": {
"currentState.$": "$.state",
"batchResults.$": "$.batchResults"
}
},
"ResultPath": "$.state",
"Next": "CheckComplete"
},
"CheckComplete": {
"Type": "Choice",
"Choices": [
{
"Variable": "$.hasMoreItems",
"BooleanEquals": true,
"Next": "GetNextBatch"
}
],
"Default": "FinaliseResults"
},
"GetNextBatch": {
"Type": "Task",
"Resource": "${GetNextBatchFunction}",
"Next": "ProcessBatch"
},
"FinaliseResults": {
"Type": "Task",
"Resource": "${FinaliseFunction}",
"End": true
}
}
}
Monitoring and Logging Best Practices
Key Concepts
Effective observability in serverless requires structured logging, distributed tracing, and custom metrics due to the ephemeral nature of function instances.
Three Pillars of Observability:
- Logs: Structured JSON logs with correlation IDs
- Metrics: Custom business and performance metrics
- Traces: Distributed tracing across services
graph TB
subgraph "Observability Pipeline"
F1[Function A] --> |"Logs"| CW[CloudWatch Logs]
F1 --> |"Traces"| XRay[X-Ray]
F1 --> |"Metrics"| Metrics[CloudWatch Metrics]
F2[Function B] --> CW
F2 --> XRay
F2 --> Metrics
CW --> Insights[CloudWatch Insights]
Metrics --> Alarms[CloudWatch Alarms]
Alarms --> SNS[Notification]
CW --> |"Export"| S3[(S3 Archive)]
S3 --> Athena[Athena Analytics]
end
Common Patterns
Structured Logging
const { Logger } = require('@aws-lambda-powertools/logger');
const logger = new Logger({
serviceName: 'order-service',
logLevel: 'INFO',
persistentLogAttributes: {
environment: process.env.ENVIRONMENT
}
});
exports.handler = async (event, context) => {
// Add correlation ID to all logs
logger.addContext(context);
logger.appendKeys({
correlationId: event.headers?.['x-correlation-id'] || context.awsRequestId
});
logger.info('Processing order request', {
orderId: event.body?.orderId,
customerId: event.body?.customerId
});
try {
const result = await processOrder(event);
logger.info('Order processed successfully', {
orderId: result.orderId,
processingTime: result.duration
});
return { statusCode: 200, body: JSON.stringify(result) };
} catch (error) {
logger.error('Order processing failed', {
error: error.message,
stack: error.stack,
orderId: event.body?.orderId
});
throw error;
}
};
Custom Metrics
const { Metrics, MetricUnits } = require('@aws-lambda-powertools/metrics');
const metrics = new Metrics({
namespace: 'OrderService',
serviceName: 'order-processor'
});
exports.handler = async (event, context) => {
// Add default dimensions
metrics.addDimension('Environment', process.env.ENVIRONMENT);
const startTime = Date.now();
try {
const result = await processOrder(event);
// Business metrics
metrics.addMetric('OrdersProcessed', MetricUnits.Count, 1);
metrics.addMetric('OrderValue', MetricUnits.None, result.total);
// Performance metrics
const duration = Date.now() - startTime;
metrics.addMetric('ProcessingDuration', MetricUnits.Milliseconds, duration);
// Cold start tracking
if (event.__coldStart) {
metrics.addMetric('ColdStarts', MetricUnits.Count, 1);
}
return result;
} catch (error) {
metrics.addMetric('OrdersFailed', MetricUnits.Count, 1);
throw error;
} finally {
// Publish all metrics
metrics.publishStoredMetrics();
}
};
Distributed Tracing
const { Tracer } = require('@aws-lambda-powertools/tracer');
const tracer = new Tracer({ serviceName: 'order-service' });
// Instrument AWS SDK clients
const { DynamoDBClient } = require('@aws-sdk/client-dynamodb');
const dynamoClient = tracer.captureAWSv3Client(new DynamoDBClient({}));
exports.handler = async (event, context) => {
// Create custom subsegment for business logic
const segment = tracer.getSegment();
const subsegment = segment.addNewSubsegment('ProcessOrder');
try {
// Add annotations for filtering traces
tracer.putAnnotation('OrderId', event.orderId);
tracer.putAnnotation('CustomerId', event.customerId);
// Add metadata for debugging
tracer.putMetadata('orderDetails', {
items: event.items,
total: event.total
});
const result = await processOrder(event);
subsegment.close();
return result;
} catch (error) {
subsegment.addError(error);
subsegment.close();
throw error;
}
};
// Decorator pattern for tracing methods
const processOrder = tracer.captureMethod({
className: 'OrderProcessor',
methodName: 'processOrder'
})(async (event) => {
// Implementation
});
Examples
CloudWatch Logs Insights Queries
-- Find slow functions
fields @timestamp, @requestId, @duration
| filter @type = "REPORT"
| filter @duration > 1000
| sort @duration desc
| limit 100
-- Error rate by function
fields @timestamp, @logStream
| filter @message like /ERROR/
| stats count() as errorCount by bin(1h)
| sort @timestamp desc
-- Cold start analysis
fields @timestamp, @duration, @billedDuration, @memorySize, @maxMemoryUsed
| filter @type = "REPORT"
| filter ispresent(@initDuration)
| stats avg(@initDuration) as avgColdStart,
count() as coldStarts
by bin(1h)
-- Trace correlation
fields @timestamp, @message, correlationId
| filter correlationId = "abc-123-def"
| sort @timestamp asc
Alarm Configuration
Resources:
ErrorAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: OrderService-HighErrorRate
AlarmDescription: Error rate exceeds 5%
MetricName: Errors
Namespace: AWS/Lambda
Dimensions:
- Name: FunctionName
Value: !Ref OrderProcessorFunction
Statistic: Sum
Period: 300
EvaluationPeriods: 2
Threshold: 5
ComparisonOperator: GreaterThanThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Ref AlertTopic
DurationAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: OrderService-HighLatency
MetricName: Duration
Namespace: AWS/Lambda
Dimensions:
- Name: FunctionName
Value: !Ref OrderProcessorFunction
ExtendedStatistic: p99
Period: 300
EvaluationPeriods: 3
Threshold: 5000
ComparisonOperator: GreaterThanThreshold
AlarmActions:
- !Ref AlertTopic
Cost Optimisation Techniques
Key Concepts
Serverless costs are based on execution duration, memory allocation, and request volume. Optimisation requires understanding the cost model and right-sizing resources.
Cost Components:
- Compute: Duration × Memory allocation
- Requests: Per-invocation charges
- Data Transfer: Cross-region and internet egress
- Storage: Logs, packages, and associated services
graph LR
subgraph "Cost Optimisation Levers"
M[Memory/CPU] --> |"Right-size"| Cost[Total Cost]
D[Duration] --> |"Optimise code"| Cost
R[Requests] --> |"Batch/cache"| Cost
T[Data Transfer] --> |"Edge caching"| Cost
end
Common Patterns
Right-Sizing Memory
// Power tuning - test different memory configurations
const memoryConfigurations = [128, 256, 512, 1024, 1536, 2048, 3008];
// Use AWS Lambda Power Tuning tool
// https://github.com/alexcasalboni/aws-lambda-power-tuning
/*
Step Functions state machine configuration:
{
"lambdaARN": "arn:aws:lambda:eu-west-1:123456789:function:my-function",
"powerValues": [128, 256, 512, 1024, 2048, 3008],
"num": 50,
"payload": {"test": "event"},
"parallelInvocation": true,
"strategy": "cost" // or "speed" or "balanced"
}
*/
// Cost calculation example
const calculateCost = (memoryMB, durationMs, requests) => {
const GB_SECOND_PRICE = 0.0000166667; // per GB-second
const REQUEST_PRICE = 0.0000002; // per request
const gbSeconds = (memoryMB / 1024) * (durationMs / 1000);
const computeCost = gbSeconds * GB_SECOND_PRICE;
const requestCost = requests * REQUEST_PRICE;
return computeCost + requestCost;
};
Batching and Aggregation
// SQS batch processing with partial failure handling
exports.handler = async (event) => {
const batchItemFailures = [];
// Process records in parallel batches
const results = await Promise.allSettled(
event.Records.map(async (record) => {
try {
await processMessage(JSON.parse(record.body));
return { success: true, messageId: record.messageId };
} catch (error) {
console.error(`Failed to process ${record.messageId}:`, error);
return { success: false, messageId: record.messageId };
}
})
);
// Report only failed items for retry
results.forEach((result, index) => {
if (result.status === 'rejected' || !result.value.success) {
batchItemFailures.push({
itemIdentifier: event.Records[index].messageId
});
}
});
return { batchItemFailures };
};
// SAM configuration for batch processing
/*
Events:
SQSEvent:
Type: SQS
Properties:
Queue: !GetAtt ProcessingQueue.Arn
BatchSize: 100
MaximumBatchingWindowInSeconds: 5
FunctionResponseTypes:
- ReportBatchItemFailures
*/
Caching Strategies
const { CacheClient, CredentialProvider, Configurations } = require('@gomomento/sdk');
// Initialise cache client outside handler
let cacheClient;
const getCache = async () => {
if (!cacheClient) {
cacheClient = await CacheClient.create({
configuration: Configurations.Lambda.latest(),
credentialProvider: CredentialProvider.fromEnvironmentVariable({
environmentVariableName: 'MOMENTO_API_KEY'
}),
defaultTtlSeconds: 300
});
}
return cacheClient;
};
exports.handler = async (event) => {
const cache = await getCache();
const cacheKey = `product:${event.productId}`;
// Try cache first
const cached = await cache.get('products', cacheKey);
if (cached.type === 'Hit') {
return {
statusCode: 200,
body: cached.valueString(),
headers: { 'X-Cache': 'HIT' }
};
}
// Fetch from database
const product = await fetchProductFromDB(event.productId);
// Cache for future requests
await cache.set('products', cacheKey, JSON.stringify(product));
return {
statusCode: 200,
body: JSON.stringify(product),
headers: { 'X-Cache': 'MISS' }
};
};
Examples
Cost Monitoring Dashboard
# CloudFormation for cost monitoring
Resources:
CostDashboard:
Type: AWS::CloudWatch::Dashboard
Properties:
DashboardName: ServerlessCostDashboard
DashboardBody: !Sub |
{
"widgets": [
{
"type": "metric",
"properties": {
"title": "Lambda Invocations & Cost",
"metrics": [
["AWS/Lambda", "Invocations", "FunctionName", "${FunctionName}"],
[".", "Duration", ".", ".", {"stat": "Sum"}],
[".", "ConcurrentExecutions", ".", "."]
],
"period": 3600,
"region": "${AWS::Region}"
}
},
{
"type": "metric",
"properties": {
"title": "Estimated Costs",
"metrics": [
[{
"expression": "(m1/1000) * (m2/1024) * 0.0000166667",
"label": "Compute Cost (USD)",
"id": "cost"
}],
["AWS/Lambda", "Duration", "FunctionName", "${FunctionName}", {"id": "m1", "stat": "Sum", "visible": false}],
[".", "Invocations", ".", ".", {"id": "m2", "stat": "Sum", "visible": false}]
],
"period": 86400
}
}
]
}
Reserved Concurrency for Cost Control
Resources:
CostControlledFunction:
Type: AWS::Serverless::Function
Properties:
FunctionName: cost-controlled-processor
Handler: index.handler
Runtime: nodejs18.x
MemorySize: 256
Timeout: 30
# Limit maximum concurrent executions
ReservedConcurrentExecutions: 100
# Use ARM architecture for cost savings (up to 34% cheaper)
Architectures:
- arm64
# Optimise deployment package size
CodeUri: ./dist
Environment:
Variables:
# Use environment variables instead of Parameter Store for static config
TABLE_NAME: !Ref DataTable
Quick Reference
| Pattern | When to Use | Key Benefits | Trade-offs |
|---|---|---|---|
| Orchestration | Complex workflows, transactions | Clear flow, centralised control | Orchestrator dependency |
| Choreography | Simple reactions, scalability | Loose coupling, scalability | Harder to trace |
| Event Sourcing | Audit trails, replay capability | Full history, debugging | Storage costs |
| CQRS | Read-heavy workloads | Optimised reads/writes | Complexity |
| Saga | Distributed transactions | Eventual consistency | Compensation logic |
| Fan-out | Parallel processing | High throughput | Cost at scale |
| Provisioned Concurrency | Latency-critical | No cold starts | Higher cost |
| Caching | Repeated reads | Lower latency, cost | Staleness |
Essential CLI Commands
# Deploy serverless application
sam deploy --guided
# Invoke function locally
sam local invoke FunctionName -e event.json
# View function logs
sam logs -n FunctionName --tail
# Power tuning analysis
aws stepfunctions start-execution \
--state-machine-arn arn:aws:states:region:account:stateMachine:powerTuningStateMachine \
--input '{"lambdaARN": "arn:aws:lambda:region:account:function:myFunction"}'
# Check function configuration
aws lambda get-function-configuration --function-name myFunction
# Update memory/timeout
aws lambda update-function-configuration \
--function-name myFunction \
--memory-size 1024 \
--timeout 30
# Enable X-Ray tracing
aws lambda update-function-configuration \
--function-name myFunction \
--tracing-config Mode=Active
# View provisioned concurrency
aws lambda get-provisioned-concurrency-config \
--function-name myFunction \
--qualifier live
Common Issues and Solutions
| Issue | Cause | Solution |
|---|---|---|
| High cold start latency | Large deployment package, VPC config | Minimise dependencies, use provisioned concurrency, SnapStart for Java |
| Timeout errors | Long-running operations | Increase timeout, use async patterns, Step Functions for workflows |
| Throttling | Concurrent execution limits | Request limit increase, implement exponential backoff, use SQS buffering |
| Out of memory | Insufficient allocation | Profile memory usage, increase allocation (also increases CPU) |
| Database connection exhaustion | Too many concurrent connections | Use RDS Proxy, connection pooling, reduce concurrency |
| Duplicate processing | At-least-once delivery | Implement idempotency with DynamoDB conditional writes |
| Lost events | No DLQ configured | Configure Dead Letter Queues on all async invocations |
| High costs | Over-provisioned memory, inefficient code | Right-size with power tuning, optimise hot paths, enable ARM |
| Debugging difficulty | Distributed architecture | Structured logging, X-Ray tracing, correlation IDs |
| State loss between invocations | Stateless functions | Externalise state to DynamoDB, Redis, or Step Functions |
Debugging Checklist
# 1. Check function logs
aws logs filter-log-events \
--log-group-name /aws/lambda/myFunction \
--filter-pattern "ERROR"
# 2. View X-Ray traces
aws xray get-trace-summaries \
--start-time $(date -d '1 hour ago' +%s) \
--end-time $(date +%s)
# 3. Check function metrics
aws cloudwatch get-metric-statistics \
--namespace AWS/Lambda \
--metric-name Errors \
--dimensions Name=FunctionName,Value=myFunction \
--start-time $(date -d '1 hour ago' -u +%Y-%m-%dT%H:%M:%SZ) \
--end-time $(date -u +%Y-%m-%dT%H:%M:%SZ) \
--period 300 \
--statistics Sum
# 4. Review DLQ messages
aws sqs receive-message \
--queue-url https://sqs.region.amazonaws.com/account/dlq-name \
--max-number-of-messages 10
# 5. Check concurrent executions
aws lambda get-account-settings
Related Topics
The following topics would complement this serverless architecture patterns cheatsheet:
-
API Gateway Patterns - Rate limiting, caching, authorisation strategies, and request/response transformations for serverless APIs
-
Infrastructure as Code (Terraform/SAM/CDK) - Defining and deploying serverless infrastructure programmatically with best practices
-
Event-Driven Messaging (SQS/SNS/EventBridge) - Deep dive into message queuing patterns, event routing rules, and guaranteed delivery
-
DynamoDB Design Patterns - Single-table design, access patterns, GSI strategies, and optimising for serverless workloads
-
Security for Serverless - IAM least privilege, secrets management, VPC configurations, and compliance considerations
-
CI/CD for Serverless - Deployment strategies (canary, blue/green), testing approaches, and pipeline automation for serverless applications