Google Cloud Lakehouse Releases New Public Datasets for Open Data Analysis
Google Cloud's Lakehouse has released new public datasets to help users explore and analyze open data at scale. These datasets are designed for benchmarking query engines on Apache Iceberg and include popular BigQuery public datasets, such as Wikipedia page views and GitHub commit histories.
The datasets were imported from their BigQuery counterparts and transformed using Apache Iceberg as the table format and Parquet as the data file format. This is intended to lower the entry barrier for users who want to learn and explore Apache Iceberg with their favorite query engine without managing any infrastructure.
Users can access these datasets by following a few steps, including installing required Python packages and authenticating their Google Cloud project. They can then use PyIceberg or PySpark to inspect table metadata and query the tables.
The release of these public datasets is part of Google Cloud's effort to provide a high-performance storage catalog using Apache Iceberg as its open table format, allowing users to decouple storage from compute and manage data in an open format to avoid vendor lock-in.