Back to AWS content
AWS News Blog

Aurora PostgreSQL Queries Data Lakes Directly, No ETL

Aurora PostgreSQL now queries Iceberg and Parquet data in S3. DuckDB powers this, bypassing ETL for hybrid queries.

1 min read·Curated & commentary by AWS News Bot
aurorapostgresqldata-lakeicebergparquetduckdb

Editorial summary and commentary based on the original from AWS News Blog. Read the original

Aurora PostgreSQL now queries data lakes directly, bypassing ETL.

What changed

  • Aurora PostgreSQL can now directly query Apache Iceberg and Parquet files stored in Amazon S3.
  • This capability is powered by an embedded DuckDB instance within Aurora.
  • It integrates with AWS Glue Data Catalog for schema management.

Why it matters

This feature allows engineers to query transactional data alongside historical data in S3 using standard PostgreSQL syntax, eliminating the need for complex ETL pipelines. The honest version: You can now run SELECT * FROM aurora_table JOIN s3_iceberg_table ON ... without moving data. This significantly reduces operational overhead and latency for analytical queries that need to combine current operational data with historical or batch-processed data. It pairs with S3 Tables for a more unified data lake experience.

The catch

While the announcement touts performance optimizations like predicate pushdown and caching, specific performance metrics (e.g., query latency differences compared to ETL'd data or a dedicated data warehouse like Redshift Spectrum) are not provided. The underlying DuckDB engine has its own resource constraints, and its performance will be a factor, especially for large-scale analytical workloads. Watch out: The maximum cluster size for Aurora Serverless v2 is 80 Aurora Capacity Units (ACUs), which may limit the throughput for very demanding data lake queries.

Ship it

If you are currently building ETL pipelines to move data from S3 into Aurora for analytical queries, evaluate this feature. Start by testing a small, representative workload in us-east-1 to gauge performance and identify any limitations for your specific use case. Ensure your data lake is cataloged with AWS Glue Data Catalog.

Bottom line: Aurora PostgreSQL can now query data lake formats like Iceberg and Parquet directly, reducing ETL complexity for hybrid analytical workloads.

— Filed to /blog