Amazon Aurora PostgreSQL Now Supports Direct Querying of Data Lakes
Amazon has introduced a new capability for its Aurora PostgreSQL database that allows direct querying of data stored in Apache Iceberg and Parquet formats in users' data lakes. This feature eliminates the need to extract, transform, and load (ETL) structured data from data lakes into the operational database, reducing operational complexity and simplifying application development.
With this capability, users can query data across external catalogs through AWS Glue Data Catalog federation without moving or duplicating it. The process is well-documented in the Aurora PostgreSQL documentation, and users can complete the setup through the Amazon RDS console or with any PostgreSQL client such as psql.
Aurora applies optimizations like predicate pushdown and column pruning to keep queries efficient even as the underlying data grows. Frequently accessed data is cached in the user's instance, so subsequent queries against the same data return faster.