AWS field guide
Keep analytical data open and queryable
Keep analytical data in an open table format so storage and compute can evolve independently.
Use this pattern when
Teams need lake-scale storage, schema evolution, and SQL access across operational and analytical tools.
Reference architecture
Responsibilities and controls, not a deployment template.
Data producers
Batch or streaming writes
Amazon S3
Durable object storage
Iceberg tables
Open table metadata
Amazon Athena
Serverless SQL
Amazon Redshift
Warehouse analytics
Decisions that shape the pattern
- Choose partition fields from real query patterns.
- Define schema ownership.
- Schedule compaction and snapshot expiry.
Security boundaries
- Control access at catalog, table, and object layers.
- Encrypt data and catalog metadata.
- Separate producer and analyst permissions.
Reliability posture
- Use atomic table commits.
- Monitor failed maintenance jobs.
- Test schema evolution with every query engine.
Starter implementation
Start from deployable infrastructure
Review every permission, limit, Region, and cost assumption before production.
import { Duration, Stack, StackProps } from 'aws-cdk-lib';
import * as apigateway from 'aws-cdk-lib/aws-apigateway';
import * as athena from 'aws-cdk-lib/aws-athena';
import * as bedrock from 'aws-cdk-lib/aws-bedrock';
import * as budgets from 'aws-cdk-lib/aws-budgets';
import * as cloudfront from 'aws-cdk-lib/aws-cloudfront';
import * as origins from 'aws-cdk-lib/aws-cloudfront-origins';
import * as cloudtrail from 'aws-cdk-lib/aws-cloudtrail';
import * as cloudwatch from 'aws-cdk-lib/aws-cloudwatch';
import * as dynamodb from 'aws-cdk-lib/aws-dynamodb';
import * as ecs from 'aws-cdk-lib/aws-ecs';
import * as patterns from 'aws-cdk-lib/aws-ecs-patterns';
import * as events from 'aws-cdk-lib/aws-events';
import * as targets from 'aws-cdk-lib/aws-events-targets';
import * as glue from 'aws-cdk-lib/aws-glue';
import * as iam from 'aws-cdk-lib/aws-iam';
import * as kms from 'aws-cdk-lib/aws-kms';
import * as lambda from 'aws-cdk-lib/aws-lambda';
import * as sources from 'aws-cdk-lib/aws-lambda-event-sources';
import * as s3 from 'aws-cdk-lib/aws-s3';
import * as secretsmanager from 'aws-cdk-lib/aws-secretsmanager';
import * as sqs from 'aws-cdk-lib/aws-sqs';
import { Construct } from 'constructs';
export class PatternStack extends Stack {
constructor(scope: Construct, id: string, props?: StackProps) {
super(scope, id, props);
const lake = new s3.Bucket(this, 'Lake', {
encryption: s3.BucketEncryption.S3_MANAGED,
blockPublicAccess: s3.BlockPublicAccess.BLOCK_ALL,
versioned: true,
});
new glue.CfnDatabase(this, 'Catalog', {
catalogId: Stack.of(this).account,
databaseInput: { name: 'open_lakehouse' },
});
new athena.CfnWorkGroup(this, 'Queries', {
name: 'governed-queries',
workGroupConfiguration: {
enforceWorkGroupConfiguration: true,
resultConfiguration: { outputLocation: 's3://' + lake.bucketName + '/athena-results/' },
},
});
}
}
Before production
Adoption checklist
- 01Profile query predicates.
- 02Set target file sizes.
- 03Automate compaction.
- 04Define retention for snapshots.
- 05Track bytes scanned per workload.
From the journal
Related field notes
Selected from service names and architecture signals used by this pattern.
Aurora PostgreSQL Queries Data Lakes Directly, No ETL
Aurora PostgreSQL now queries Iceberg and Parquet data in S3. DuckDB powers this, bypassing ETL for hybrid queries.
Redshift Cross-Region Data Lake Queries Now GA
Query S3 data lakes in other regions directly from Redshift. Finally.
S3 Tables Add Full Apache Iceberg V3 Support
S3 Tables now align with the full Apache Iceberg V3 specification, bringing enhanced data type compatibility and features.
Was this playbook useful?
One click helps prioritize deeper examples and updates.