1 New public datasets available in Google Cloud Lakehouse | Google Open Source Blog

opensource.google.com

Menu

New public datasets available in Google Cloud Lakehouse

Wednesday, September 30, 2026

An image of a castle with an ice block coming out of the front gate

Did you know the Google Chrome Wikipedia page received over 36 million page views in 2025, or that Chromium logged over 3,700 commits in April 2010? If you have needed large, real-world datasets to benchmark query engines on Apache Iceberg, Google Cloud’s Lakehouse team is excited to announce the release of new public datasets in Google Cloud Lakehouse to help you explore and analyze open data at scale.

Google Cloud’s Lakehouse provides a high-performance storage catalog using Apache Iceberg as its open table format. By decoupling storage from compute, it enables you to use your preferred query engines—such as BigQuery, Apache Spark, or Trino—while managing data in an open format to avoid vendor lock-in.

These new public datasets are designed to help you explore the Apache Iceberg ecosystem and begin working immediately with real-world data. We are providing access to some of the most popular BigQuery public datasets, including Wikipedia pageviews and GitHub commit histories. Let’s dive into how you can start querying them.

Prerequisites

  • A Google Cloud project (for authentication).
  • Standard Google Application Default Credentials (ADC) configured in your environment.

Explore metadata with PyIceberg

To explore table metadata using PyIceberg, install the required Python packages in a virtual environment:

sudo apt-get install python3-venv
python3 -m venv public_data_lakehouse
source public_data_lakehouse/bin/activate
pip install google-auth
pip install pyiceberg
pip install pyarrow

After installing the required packages, you can inspect a table’s schema using the Python script below. Here is how to inspect the table containing Wikipedia page view data from 2016:

from pyiceberg import catalog as pyiceberg_catalog

PROJECT = "<YOUR_PROJECT_ID>"
NAMESPACE = "wikipedia"
TABLE = "pageviews_2016"

def load_catalog() -> pyiceberg_catalog.Catalog:
  properties = {
      "type": "rest",
      "uri": "https://biglake.googleapis.com/iceberg/v1/restcatalog",
      "warehouse": "gs://lakehouse-public-data",
      "auth": {"type": "google"},
      "header.X-Iceberg-Access-Delegation": "vended-credentials",
      "header.x-goog-user-project": PROJECT,
  }
  return pyiceberg_catalog.load_catalog(PROJECT, **properties)

def main() -> None:
  catalog = load_catalog()
  table = catalog.load_table((NAMESPACE, TABLE))
  print(f"Schema for `{NAMESPACE}.{TABLE}`:", table.schema())

if __name__ == "__main__":
  main()

Note: Replace <YOUR_PROJECT_ID> with your actual Google Cloud Project ID. This is required for the REST catalog to authenticate your quota usage, even for free public access.

PySpark and Managed Service for Apache Spark

Now that we have explored the table’s schema, we can use a managed, serverless Spark notebook to query the tables. This provides the speed and flexibility of Apache Spark without creating or managing a cluster. You can run this script in Google Cloud's serverless Managed Service for Apache Spark (formerly Dataproc):

from google.cloud.dataproc_v1 import Session
from google.cloud.dataproc_spark_connect import DataprocSparkSession

PROJECT_ID = "<YOUR_PROJECT_ID>"
spark_catalog = "lakehouse-public-data"

session = Session()
session.runtime_config.properties = {
  "spark.sql.extensions": "org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions",
  f"spark.sql.catalog.{spark_catalog}": "org.apache.iceberg.spark.SparkCatalog",
  f"spark.sql.catalog.{spark_catalog}.type": "rest",
  f"spark.sql.catalog.{spark_catalog}.uri": "https://biglake.googleapis.com/iceberg/v1/restcatalog",
  f"spark.sql.catalog.{spark_catalog}.rest.auth.type": "org.apache.iceberg.gcp.auth.GoogleAuthManager",
  f"spark.sql.catalog.{spark_catalog}.io-impl": "org.apache.iceberg.gcp.gcs.GCSFileIO",
  f"spark.sql.catalog.{spark_catalog}.header.X-Iceberg-Access-Delegation": "vended-credentials",
  f"spark.sql.catalog.{spark_catalog}.warehouse": "gs://lakehouse-public-data",
  f"spark.sql.catalog.{spark_catalog}.header.x-goog-user-project": PROJECT_ID,
}

spark = (
   DataprocSparkSession.builder
     .appName("Lakehouse Public Data Demo")
     .dataprocSessionConfig(session)
     .getOrCreate()
)
spark.conf.set("spark.sql.defaultCatalog", spark_catalog)

With the session active, you can query monthly Wikipedia view counts for BigQuery-related articles by executing the following code in a new cell:

df1 = spark.sql("""
SELECT
  title,
  wiki,
  SUM(views) AS total_views
FROM wikipedia.pageviews_2026
WHERE datehour >= TIMESTAMP '2026-01-01 00:00:00'
  AND datehour <  TIMESTAMP '2026-02-01 00:00:00'
  AND LOWER(title) LIKE '%bigquery%'
GROUP BY title, wiki
ORDER BY total_views DESC
LIMIT 20
""")
df1.show(10)

Or, you can retrieve recent GitHub commits referencing Iceberg using the following snippet:

df2 = spark.sql("""
SELECT
  commits.commit,
  commits.subject,
  commits.message,
  commits.author.name AS author_name,
  timestamp_seconds(commits.committer.date.seconds) AS commit_time,
  repo_name
FROM github_repos.commits AS commits
WHERE LOWER(commits.subject) LIKE '%iceberg%'
   OR LOWER(commits.message) LIKE '%iceberg%'
ORDER BY commit_time DESC
LIMIT 50
""")
df2.show(10)

Start building today

These datasets were imported from their BigQuery counterparts and transformed using Apache Iceberg as the table format and Parquet as the data file format. Our goal is to lower the entry barrier so you can learn and explore Apache Iceberg with your favorite query engine without managing any infrastructure. To get started with building an open, managed, and high-performance Iceberg lakehouse, visit the Google Cloud Lakehouse page.

.