opensource.google.com

Menu

Standardizing fine-grained access control in Apache Iceberg's REST catalog

Thursday, October 1, 2026

The Apache Iceberg community recently merged a change to the REST Catalog specification that gives lakehouses something they have long been missing: a standard, engine-neutral way to express column masking and row filtering. It is a small change, but it addresses a structural gap in how open table formats handle data governance. This post walks through the problem, the design, and why the details matter.

Finally a single fine-grained access control standard for any engine

The problem: there is no server in the read path

In a traditional database, access control is straightforward because there is exactly one door. Every query passes through the database server, the server knows who is asking, and if policy says "this user only sees the last four digits of the card number," the server applies that policy before returning results. One process, one enforcement point.

Apache Iceberg deliberately removed that door. The data is Parquet files in object storage, and any engine—such as BigQuery, Spark, Trino, Flink, PyIceberg, or DuckDB—can read those files directly. This design is what makes Iceberg fast and interoperable. But it also means that once an engine asks the catalog "where is the payments table?" and receives the metadata pointer, it has full access to the files. The catalog can grant or deny access to the whole table, but it has had no vocabulary for "access granted, but mask the email column" or "access granted, but only on rows where region = 'US'."

Until now, Iceberg had no built-in answer to this. The REST spec did offer coarser instruments like credential vending, which controls storage access per table, and server-side scan planning, which lets a catalog withhold entire data files. But neither can express "mask this column" or filter rows that are not already physically separated into their own files. So users who needed fine-grained access control had to step outside the open protocol entirely. They could adopt a vendor-provided client that understands its proprietary policy format, or route every read through a vendor's proxy service, giving up the direct-to-storage performance that motivated the use of Iceberg in the first place. It also quietly undermines Iceberg's core promise: the moment governance requires a specific vendor's client, the table is no longer open to any engine.

The new read-restrictions field in the REST Catalog spec is the community standardizing that vocabulary.

The mechanism

When an engine calls loadTable, the response may now include an optional read-restrictions object:

Read restrictions payload showing required column projections and required row projections

Read restrictions are expressed in two fields:

  • required-row-filter: a standard Iceberg predicate expression. Rows for which it evaluates to false must not appear in the result, and no information derived from them may be included.
  • required-column-projections: a list of columns, identified by field ID, each with a transformation the reader must apply before returning values.

One evaluation rule ties them together: the row filter is evaluated against the original, untransformed column values, and projections are applied to the rows that survive. This ordering is what makes the two features composable. A policy can filter on region = 'US' and also mask region in the output, and the filter still works. If masking ran first, any policy that filtered on a masked column would silently break.

Note what the catalog is doing here. It evaluates the access policy server-side—it knows the caller's identity from the authentication token—and returns only the result of that evaluation. The policy itself, with its roles, tags, and governance model, never crosses the wire. The engine does not need to understand how any particular catalog models governance. It needs to understand nine actions and a predicate.

The whole architecture fits in one sequence:

Sequence diagram illustrating the Lakehouse credential vending and reader-side fine-grained access control workflow between an End User, a Trusted Engine (BigQuery/Spark), the Lakehouse runtime catalog, and Google Cloud Storage (GCS).
The catalog stays the policy decision point; the trusted engine becomes the policy enforcement point; the end user never holds storage credentials. Sections below unpack the three load-bearing details in this picture: the enforcement ordering, the fail-closed branch, and the trust boundary.

The nine masking actions

The specification defines a closed set of nine masking actions, each with an exact, per-type definition. The goal is cross-engine consistency: Spark, Trino, and PyIceberg must produce identical output for any given masking action. These actions are being implemented in iceberg-core, providing engines with spec-compliant transformations out of the box rather than requiring them to reimplement byte-level logic independently.

Below, the actions are grouped by the analytical utility of the resulting masked data:

Preserve the shape, hide the value

  • mask-alphanum—digits become n, other characters become x, with a small allowlist of punctuation (( ) , . - @) preserved. iceberg16112018@apache.org becomes xxxxxxxnnnnnnnn@xxxxxx.xxx—recognizably an email address, but not whose.
  • show-first-4 / show-last-4—preserve four code points, and apply mask-alphanum to the rest. 4111-1111-1111-4444 becomes nnnn-nnnn-nnnn-4444, the familiar customer-support view of a card number.

Hide everything

  • replace-with-null—the value becomes NULL. Only valid for optional fields; a server must not return it for a required field, and a reader that receives one must fail the query.
  • mask-to-fixed-value—the value becomes a type-specific constant (0, "XXXXXXXX", the epoch, an all-zero UUID, an empty list, and so on, each spelled out in the spec). Uniquely among the actions, this one also replaces NULL inputs—even the null-or-not bit is hidden.

Reduce precision

  • truncate-to-year / truncate-to-month—2024-07-15 becomes 2024-01-01 or 2024-07-01. Cohort analytics keeps working; identifying individuals gets harder.

Allow joins, hide values

  • sha-256-global—deterministic SHA-256, with exact byte-encoding rules per input type. The same input always produces the same output, everywhere, so GROUP BY user_id and joins across tables on a hashed key still work. The cost of that determinism is that hashed values can be tested against precomputed guesses—this is pseudonymization, not encryption.
  • sha-256-query-local—the same hash, salted with a fresh, cryptographically random salt (at least 16 bytes) per query. Values remain consistent within a single query, so self-joins and aggregation work, but cannot be correlated across queries, and precomputed-guess attacks no longer apply.

This last pair is a nice piece of design: the tradeoff between linkability and privacy, expressed as two enum values that a policy author chooses between per column.

Every action produces a value of the same type as its input—masked strings are strings, truncated dates are dates—so restrictions never change the schema an engine plans against. Queries do not need rewriting; values simply arrive transformed.

Fail-closed by design

The most consequential sentence in the specification is this one:

If a trusted reader that supports read-restrictions cannot apply any returned restriction, it must fail the query and must not silently return raw, partial, or empty results.

Consider the failure modes. A catalog sends an action added in a future spec version that the engine does not recognize. Or an expression type it cannot evaluate. The convenient behavior would be to skip what it does not understand and return the data. The spec rules this out: unrecognized action, fail; unparseable filter, fail; duplicate field ID in the projections, fail. Every ambiguity resolves to "no data" rather than "raw data." This is the right default for an access control mechanism, and it is also the one implementers would be tempted to soften—which is exactly why it is normative in the spec rather than left to judgment. It is also what allows the action vocabulary to grow in future versions without older engines becoming silent leak vectors.

A few prohibitions in the spec reward a closer look, because each one closes a subtle correctness hole:

  • No projections on map keys. Masking a map's keys can collapse two keys into one, or produce null keys, which engines silently coalesce or reject—data corruption presented as privacy. The spec bans it outright.
  • No projection on both a nested type and a field inside it. Masking a struct and also a field within that struct has no well-defined order of operations, so the spec refuses to define one: servers must not send it, and readers must reject it.
  • Everything references field IDs, never column names. This is standard Iceberg discipline: if ssn is renamed to national_id, the policy remains bound to the same physical column. A name-based policy would silently detach on rename—the worst possible failure mode for access control.

The trust model, stated plainly

All of this is enforced by the reader. The catalog hands the engine the file locations along with the restrictions, and a client that chooses to ignore the restrictions can read the raw files. So what is this actually protecting?

The specification is explicit: this mechanism assumes a trust relationship between the catalog and the engine, and how that trust is established is deliberately out of scope. The intended deployment is one where a platform team's engines—the shared Spark and Trino clusters—are trusted enforcement points that hold storage access (for example, through credential vending), while end users only ever talk to those engines and never hold storage credentials themselves. The trusted engine becomes the enforcement point, playing the role the database server played in the traditional architecture. The difference is that its behavior is now defined by a common, open protocol rather than by N proprietary integrations.

In other words, read restrictions do not protect data from the engine; they let the catalog direct a trusted engine to protect data from the engine's users. For genuinely untrusted readers, the coarse-grained model still applies: they are restricted to vended credentials where the storage access granted to the user aligns with the data permission of the user.

One operational subtlety deserves attention: restrictions are per-response and per-identity. The same loadTable call made by a different principal—or by the same principal later—may return different restrictions, and the restrictions attach to every read performed with that response, including subsequent planTableScan and fetchScanTasks calls. The spec therefore requires that the response not be cached outside its authentication scope. If your platform caches loadTableResponse, that cache is now security-sensitive and worth an audit.

What is still missing

Read restrictions are a foundation, not the finished building. It is worth being honest about the gaps between this specification and complete fine-grained access control:

  1. The action vocabulary is fixed, because Iceberg Expressions are not implemented yet. Nine actions cover the common masking policies, but they are a deliberately closed set: a policy like "apply this custom redaction function" cannot be expressed, and the row filter is limited to predicates—comparisons that produce true or false. The path to generalizing this already exists on paper: the Iceberg Expressions specification, proposed by Ryan Blue and adopted in mid-2026, defines a portable structure for value expressions—constants, field references, and calls to well-defined functions or SQL UDFs. Once engines can evaluate those expressions, a catalog could return arbitrary transformations instead of choosing from an enum. Today no engine implements general expression evaluation, which is why the initial design confines itself to a small vocabulary that can be specified bit-for-bit.
  2. There is no policy definition—deliberately. The specification standardizes the result of policy evaluation, never the policy itself. How an organization expresses "analysts see masked PII, auditors see everything"—the roles, tags, rules, and administrative APIs—remains entirely the catalog vendor's domain, and the assumption is that it stays there. Only the consequences of a policy are portable across engines; the policy definition is not. Whether communities eventually want a portable policy format is an open question the spec does not attempt to answer.
  3. Trusted clients are asserted, not proven. As the trust-model section noted, how a catalog establishes that a caller is a trusted, enforcing engine is out of scope. In practice that trust is deployment configuration—service identities, network boundaries, and which principals receive vended credentials. There is no attestation mechanism in the protocol by which an engine proves it enforces restrictions; the trusted-client mechanism is an assumption the platform operator must make true.

None of these gaps undermines the design—each is a deliberate scoping decision that kept the proposal small enough to reach consensus—but they define the roadmap for what "complete" fine-grained access control in the open lakehouse still requires.

Why this matters

The specification change itself does not ship enforcement; that work in the engines begins now, starting with the default actions. But the shape of the design is right in three ways:

  1. It picks the honest enforcement point. In an architecture with no server in the read path, the trusted engine is the only place enforcement can live without giving up direct storage reads. The spec accepts that constraint explicitly rather than obscuring it.
  2. It standardizes the narrow waist. Catalogs keep their own rich policy engines—roles, tags, and attribute-based rules. Engines implement nine actions and a predicate evaluator, once. N×M becomes N+M.
  3. It is fail-closed everywhere, which is what makes the vocabulary safely extensible.

The pattern—the catalog evaluates policy against the caller's identity and returns a small, closed vocabulary of obligations that the client must enforce or fail—is a useful template, and it would not be surprising to see more of the governance surface expressed this way over time.

The proposal was approved on the Apache Iceberg dev list with eight binding +1 votes and no objections. This is a strong signal of consensus across the many companies and open source communities that participate in the project. Consistent, engine-independent enforcement of fine-grained policies is a property the ecosystem has long wanted; it now has a specification for it, and the interesting work of implementing it in engines and catalogs is underway. If you work on either, the dev list is the place to get involved.

The full schema is available in rest-catalog-open-api.yaml under ReadRestrictions. View the full vote thread.

New public datasets available in Google Cloud Lakehouse

Wednesday, September 30, 2026

An image of a castle with an ice block coming out of the front gate

Did you know the Google Chrome Wikipedia page received over 36 million page views in 2025, or that Chromium logged over 3,700 commits in April 2010? If you have needed large, real-world datasets to benchmark query engines on Apache Iceberg, Google Cloud’s Lakehouse team is excited to announce the release of new public datasets in Google Cloud Lakehouse to help you explore and analyze open data at scale.

Google Cloud’s Lakehouse provides a high-performance storage catalog using Apache Iceberg as its open table format. By decoupling storage from compute, it enables you to use your preferred query engines—such as BigQuery, Apache Spark, or Trino—while managing data in an open format to avoid vendor lock-in.

These new public datasets are designed to help you explore the Apache Iceberg ecosystem and begin working immediately with real-world data. We are providing access to some of the most popular BigQuery public datasets, including Wikipedia pageviews and GitHub commit histories. Let’s dive into how you can start querying them.

Prerequisites

  • A Google Cloud project (for authentication).
  • Standard Google Application Default Credentials (ADC) configured in your environment.

Explore metadata with PyIceberg

To explore table metadata using PyIceberg, install the required Python packages in a virtual environment:

sudo apt-get install python3-venv
python3 -m venv public_data_lakehouse
source public_data_lakehouse/bin/activate
pip install google-auth
pip install pyiceberg
pip install pyarrow

After installing the required packages, you can inspect a table’s schema using the Python script below. Here is how to inspect the table containing Wikipedia page view data from 2016:

from pyiceberg import catalog as pyiceberg_catalog

PROJECT = "<YOUR_PROJECT_ID>"
NAMESPACE = "wikipedia"
TABLE = "pageviews_2016"

def load_catalog() -> pyiceberg_catalog.Catalog:
  properties = {
      "type": "rest",
      "uri": "https://biglake.googleapis.com/iceberg/v1/restcatalog",
      "warehouse": "gs://lakehouse-public-data",
      "auth": {"type": "google"},
      "header.X-Iceberg-Access-Delegation": "vended-credentials",
      "header.x-goog-user-project": PROJECT,
  }
  return pyiceberg_catalog.load_catalog(PROJECT, **properties)

def main() -> None:
  catalog = load_catalog()
  table = catalog.load_table((NAMESPACE, TABLE))
  print(f"Schema for `{NAMESPACE}.{TABLE}`:", table.schema())

if __name__ == "__main__":
  main()

Note: Replace <YOUR_PROJECT_ID> with your actual Google Cloud Project ID. This is required for the REST catalog to authenticate your quota usage, even for free public access.

PySpark and Managed Service for Apache Spark

Now that we have explored the table’s schema, we can use a managed, serverless Spark notebook to query the tables. This provides the speed and flexibility of Apache Spark without creating or managing a cluster. You can run this script in Google Cloud's serverless Managed Service for Apache Spark (formerly Dataproc):

from google.cloud.dataproc_v1 import Session
from google.cloud.dataproc_spark_connect import DataprocSparkSession

PROJECT_ID = "<YOUR_PROJECT_ID>"
spark_catalog = "lakehouse-public-data"

session = Session()
session.runtime_config.properties = {
  "spark.sql.extensions": "org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions",
  f"spark.sql.catalog.{spark_catalog}": "org.apache.iceberg.spark.SparkCatalog",
  f"spark.sql.catalog.{spark_catalog}.type": "rest",
  f"spark.sql.catalog.{spark_catalog}.uri": "https://biglake.googleapis.com/iceberg/v1/restcatalog",
  f"spark.sql.catalog.{spark_catalog}.rest.auth.type": "org.apache.iceberg.gcp.auth.GoogleAuthManager",
  f"spark.sql.catalog.{spark_catalog}.io-impl": "org.apache.iceberg.gcp.gcs.GCSFileIO",
  f"spark.sql.catalog.{spark_catalog}.header.X-Iceberg-Access-Delegation": "vended-credentials",
  f"spark.sql.catalog.{spark_catalog}.warehouse": "gs://lakehouse-public-data",
  f"spark.sql.catalog.{spark_catalog}.header.x-goog-user-project": PROJECT_ID,
}

spark = (
   DataprocSparkSession.builder
     .appName("Lakehouse Public Data Demo")
     .dataprocSessionConfig(session)
     .getOrCreate()
)
spark.conf.set("spark.sql.defaultCatalog", spark_catalog)

With the session active, you can query monthly Wikipedia view counts for BigQuery-related articles by executing the following code in a new cell:

df1 = spark.sql("""
SELECT
  title,
  wiki,
  SUM(views) AS total_views
FROM wikipedia.pageviews_2026
WHERE datehour >= TIMESTAMP '2026-01-01 00:00:00'
  AND datehour <  TIMESTAMP '2026-02-01 00:00:00'
  AND LOWER(title) LIKE '%bigquery%'
GROUP BY title, wiki
ORDER BY total_views DESC
LIMIT 20
""")
df1.show(10)

Or, you can retrieve recent GitHub commits referencing Iceberg using the following snippet:

df2 = spark.sql("""
SELECT
  commits.commit,
  commits.subject,
  commits.message,
  commits.author.name AS author_name,
  timestamp_seconds(commits.committer.date.seconds) AS commit_time,
  repo_name
FROM github_repos.commits AS commits
WHERE LOWER(commits.subject) LIKE '%iceberg%'
   OR LOWER(commits.message) LIKE '%iceberg%'
ORDER BY commit_time DESC
LIMIT 50
""")
df2.show(10)

Start building today

These datasets were imported from their BigQuery counterparts and transformed using Apache Iceberg as the table format and Parquet as the data file format. Our goal is to lower the entry barrier so you can learn and explore Apache Iceberg with your favorite query engine without managing any infrastructure. To get started with building an open, managed, and high-performance Iceberg lakehouse, visit the Google Cloud Lakehouse page.

This Week in Open Source for September 24, 2026

Thursday, September 24, 2026

This Week in Open Source banner graphic

This Week in Open Source for September 24, 2026

A look around the world of open source

Open source underpins an $8.8 trillion global economy, yet sustaining it requires moving beyond voluntary charity and philosophical appeals. This week’s reads examine how the ecosystem is adapting to modern economic and AI pressures—from proposing package-registry royalties and quantifying 2–5x net value for regulated utility grids, to turning autonomous AI security incidents into hardened supply-chain defenses that empower open-science breakthroughs like NASA and IBM’s Lunar Foundation Model.

Upcoming Events

🗓️ October 2026

  • MCP Dev Summit Toronto (October 5–6, 2026) — Toronto, Ontario, Canada. Dedicated Linux Foundation developer summit advancing the Model Context Protocol (MCP) and open agentic interoperability standards. Join Google OSPO's Daryl Ducharme on October 6 for "Tag-Team Transmission: Navigating A2A and MCP for Optimum Orchestration," examining how the Agent2Agent (A2A) protocol and MCP interoperate across multi-agent architectures.
  • Open Source Summit Europe (October 7–9, 2026) — Prague, Czechia. The premier Linux Foundation gathering in Europe celebrating the 35th anniversary of Linux and connecting developers, technologists, and community leaders across open AI, embedded systems, and digital trust.
  • Community Over Code (October 11–14, 2026) — Glasgow, Scotland. The flagship Apache Software Foundation conference bringing together project maintainers, committers, and users to collaborate on open governance, data architecture, and community-led software development.
  • All Things Open 2026 (October 19–20, 2026) — Raleigh, North Carolina, USA. One of the largest community-focused open source conferences on the U.S. East Coast exploring open source software, AI engineering, DevSecOps, and maintainer sustainability.
  • GitHub Universe 2026 (October 28–29, 2026) — San Francisco, California, USA & Virtual. Annual global developer gathering highlighting open source workflows, collaborative security, and AI-assisted software engineering.

🗓️ November 2026

  • Open Source in Finance Forum (OSFF) New York (November 4–5, 2026) — New York, New York, USA. Dedicated industry conference examining open source compliance, supply-chain security, and collaborative innovation across regulated financial institutions.
  • KubeCon + CloudNativeCon North America 2026 (November 9–12, 2026) — Salt Lake City, Utah, USA. The Cloud Native Computing Foundation's flagship conference uniting adopters and maintainers around Kubernetes, platform engineering, and modern cloud infrastructure.
  • SFSCON (South Tyrol Free Software Conference) 2026 (November 13–14, 2026) — Bolzano, Italy. One of Europe's longest-running Free Software conferences bringing together public-sector decision-makers and developers to advance digital sovereignty and open infrastructure.

Open Source Reads and Links

  • [Blog] Nobody pays for open source. We can force them to. — Companies already spend over $1 billion a year on open source, but that money goes to supply-chain mirror and security vendors rather than the maintainers writing the code. Laurie Voss breaks down why voluntary charity and license changes always fail, arguing that commercial package registries and mirror vendors should pay automatic royalties down the dependency tree to long-tail maintainers.
  • [Article] AI Is reshaping open source software and straining the systems that sustain it — As conversations grow around moderating the speed of AI development across the industry, it is critical to watch the friction points where AI and open source overlap—particularly as an influx of AI-generated code and vulnerability discovery strains maintainers of an $8.8 trillion ecosystem. Funding remains front and center, requiring not just raw capital but continuous monitoring for effectiveness in sustaining the human governance, consensus-building, and non-coding infrastructure that AI cannot replace.
  • [Report] LF Energy Research Finds Open Source Software Can Deliver 2-5x Greater Net Value for Grid Operators — While open source software and public utilities seem like natural philosophical allies as public goods, philosophical appeals routinely fail in regulated infrastructure sectors where operators are bound by strict reliability mandates and ratepayer accountability. This new LF Energy report demonstrates how to bridge that gap: by translating collaborative "make together" governance into an auditable benefit-cost framework showing two to five times greater net value over proprietary procurement.
  • [Article] AI Is at a Turning Point — High-profile incidents of autonomous AI agents breaching Hugging Face or social-engineering open source maintainers inevitably dominate headlines and amplify purely negative narratives around AI. However, treating these security failures as concrete post-mortems—exposing how standard reinforcement learning incentivizes deceptive subgoals—is exactly what allows open source communities to establish hardened supply-chain defenses and safely advance high-impact, positive AI use cases.
  • [Article] IBM and NASA Release Open-Source AI Model to Support Lunar Exploration — Seeing open source, AI, open data, and space science converge is genuinely inspiring, especially when it demonstrates that open scientific models are only as useful as the open datasets beneath them. While NASA and JAXA planetary archives have long been public, IBM and NASA's release of the Lunar Foundation Model alongside the first harmonized, machine-learning-ready lunar dataset (unifying 30+ multi-instrument layers) proves that transforming raw open data into shared, interoperable infrastructure is what unlocks domain discovery.

Which of these stories will you be chatting about at your next meetup or conference? Let us know! Share with us on our @GoogleOSS X account or our @opensource.google Bluesky account.

Google Cloud: Investing in the future of PostgreSQL — 2026 highlights

Wednesday, September 23, 2026

Google Cloud is committed to open source, and PostgreSQL is a cornerstone of managed database offerings, including Cloud SQL and AlloyDB.

Continuing our work with the PostgreSQL communities, we've been contributing to the core engine and participating in the patch review process. Below is a summary of that technical activity between January 2026 and September 2026, highlighting our efforts to enhance the performance, stability, and resilience of the upstream project and ecosystem. By strengthening these core capabilities, we aim to drive innovation that benefits the entire global PostgreSQL ecosystem and its diverse user base.

Our technical contributions in this period have focused on enhancing core engine performance, introducing features for logical replication, fixing critical bugs, and improving upgrade resilience. We also continue to invest in the PostgreSQL ecosystem by addressing bugs in widely used extensions.

Technical contributions: January 2026 – September 2026

Our contributions this cycle span four key areas:

  1. Logical replication and conflict management: Paving the way for active-active replication, enhanced conflict logging, and catalog safety.
  2. Core engine performance and reliability: Eliminating tuple-level overhead in sequential scans and making promotion timeouts accurate.
  3. Core catalog, collation, and indexing bug fixes: Strengthening privilege consistency, deferrable index builds, collation handling, and memory safety.
  4. PostgreSQL extension and ecosystem hardening: Resolving critical crashes, lock tranche registrations, and memory vulnerabilities across popular ecosystem extensions (plpgsql_check, pgfincore, and pgtt).

1. Logical replication and conflict management

Logical replication is essential for near-zero downtime migrations, major version upgrades, and multi-region data distribution. Our recent work focuses on conflict log table infrastructure, namespace clarity, and cross-database catalog hygiene.

Conflict log table infrastructure

  • Background and challenge: A major milestone on the roadmap to active-active multi-master replication is the ability to automatically record and resolve data discrepancies across nodes. Without structured conflict logs, tracking conflicting writes requires inspecting server logs or relying on custom application handlers.
  • Solution: This patch introduces the foundational conflict log table management infrastructure as an option in CREATE SUBSCRIPTION. It establishes the catalog structures and configuration hooks necessary to direct conflict records into dedicated, queryable tables, serving as the cornerstone for upcoming automatic conflict logging into tables.
  • Note: Provided no blocking issues arise, this functionality is slated for inclusion in PG 20.
  • Contributors: Dilip Kumar (Primary Author)

Fix REASSIGN OWNED for subscriptions in other databases

  • Background and challenge: While pg_subscription is physically a shared catalog so the background launcher process can scan all databases, subscription objects are logically local to each database. Certain operations, such as REASSIGN OWNED, failed to restrict their catalog scans to the current database (MyDatabaseId), leading to accidental modifications across database boundaries.
  • Solution: Added explicit guards to pg_subscription readers to ensure non-launcher processes filter strictly by MyDatabaseId, protecting cross-database isolation and updating documentation.
  • Contributors: Dilip Kumar (Author)

Schema-qualified names in EXCEPT clause error messages

  • Background and challenge: When publishing tables with EXCEPT clauses, check_publication_add_relation() previously reported only unqualified table names when a relation could not be processed, leading to ambiguous error messages in multi-schema databases.
  • Solution: Updated error reporting paths to emit fully schema-qualified relation names, aligning with PostgreSQL's broader error messaging standards.
  • Contributors: Dilip Kumar (Author)

2. Core engine performance and administrative enhancements

Optimizing throughput and improving the predictability of administrative operations remain top priorities for database workloads.

Timeout handling in pg_promote()

  • Background and challenge: Standby promotion via pg_promote() allows users to specify a timeout interval. Due to imprecise elapsed-time tracking during wait loops, promotion operations could terminate prematurely before the configured timeout elapsed.
  • Solution: Refined the promotion loop to pre-calculate the expected end timestamp and track actual elapsed time across iterations, ensuring strict adherence to user-configured wait intervals.
  • Contributors: Robert Pang (Reporter and Author)

Sequential scan performance optimization

  • Background and challenge: Previously, CheckXidAlive validation was executed inside the inner table_scan_next routines. This incurred repetitive check overhead on every single tuple fetched during sequential scans.
  • Solution: Restructured the scan control flow to eliminate redundant per-tuple checks, yielding cleaner execution paths and measurable throughput improvements on large table scans.
  • Contributors: Dilip Kumar (Author)

3. Core engine access control, collation, and indexing bug fixes

We continue to harden PostgreSQL's core engine against catalog inconsistencies, segmentation faults, and edge-case query anomalies.

Large object access with pg_{read,write}_all_data

  • Problem: The default roles pg_read_all_data and pg_write_all_data were designed to allow maintenance utilities like pg_dump to operate without superuser privileges. However, Large Objects (LOBs) remained inaccessible under these roles without explicit object-level grants.
  • Fix: Updated permission checks to extend pg_read_all_data and pg_write_all_data coverage to Large Objects, completing superuser-free dump and maintenance workflows.
  • Contributors: Nitin Motiani (Author), Dilip Kumar (Reviewer)

Immediate property propagation in index copies (REINDEX CONCURRENTLY)

  • Problem: When building a replacement index during REINDEX CONCURRENTLY for a deferrable unique constraint, index_create_copy() defaulted constraint flags to 0, setting the immediate property to true. This caused concurrent transactions to immediately trigger constraint violations rather than deferring verification until commit time.
  • Fix: Introduced the INDEX_CREATE_DEFERRABLE flag to properly propagate an immediate property of false to transient copied indexes without violating internal constraint assertions.
  • Contributors: Nitin Motiani (Author)

LIKE matching with nondeterministic collations and backslashes

  • Problem: Following the addition of nondeterministic collation support for LIKE, literal pattern substring parsing unconditionally skipped all backslash characters. When encountering escaped backslashes (\\), the engine omitted the second backslash entirely instead of emitting a literal \.
  • Fix: Corrected pattern de-escaping logic to correctly recognize and emit escaped backslashes during evaluation.
  • Contributors: Nitin Motiani (Author)

DSM lock release and crash prevention

  • Problem: If a backend encountered a FATAL exit while holding a lock in a Dynamic Shared Memory (DSM) segment (e.g., inside dynamic shared hashtables dshash) outside of an active transaction, releasing locks during process termination could reference already detached DSM segments, triggering a segmentation fault.
  • Fix: Hardened cleanup and lock release sequences during fatal exits to safely detach memory segments without segfaulting.
  • Contributors: Dilip Kumar (Reviewer)

4. PostgreSQL ecosystem and extension hardening

Enterprise PostgreSQL architectures rely heavily on third-party extensions. Our team actively contributes bug fixes and stability improvements upstream to critical ecosystem projects.

plpgsql_check: LWLock tranche registration for PG14

  • Problem: In PG14 and earlier, loading plpgsql_check via shared_preload_libraries could fail with shared memory lock errors due to missing or outdated named LWLock tranche registrations (plpgsql_check profiler funcs stats and plpgsql_check profiler func stmts stats).
  • Fix: Aligned pre-PG15 tranche initialization with modern shmem_request_hook patterns, guaranteeing safe shared memory allocation on older server versions.
  • Contributors: Aniket Jha (Author)

pgfincore: Memory safety hardening

  • Problem: pgfincore contained two subtle memory corruption issues: an off-by-one array boundary access during buffer inspection and a dangling pointer assignment during deallocation.
  • Fix: Authored patches to enforce strict boundary checks and clean pointer resets, eliminating potential memory corruption during OS buffer cache analysis.
  • PRs: klando/pgfincore#12 and klando/pgfincore#13
  • Contributors: Robert Pang (Author)

pgtt: Use-after-free prevention on cached plans

  • Problem: When running utility commands via the extended query protocol or within cached contexts (such as PL/pgSQL and SQL functions), pgtt modified cached statement parse trees in-place using short-lived query memory. Once that memory was freed, subsequent executions of the cached plan led to Use-After-Free crashes.
  • Fix: Updated the extension to operate on an isolated, deep copy of the parse tree for cached and read-only statements, ensuring memory safety across repeated executions.
  • Contributors: Sunaina Punyani (Author)

Related reading

Community roadmap: Your feedback matters

We encourage you to utilize the comments area to propose new capabilities or refinements you wish to see in future iterations, and to identify key areas where the PostgreSQL open source communities should focus their investments.

Acknowledgments

We would like to celebrate our engineers for their ongoing dedication to open source:

  • Dilip Kumar (PostgreSQL Significant Contributor): Authoring and reviewing core replication, catalog, memory, and performance patches.
  • Nitin Motiani: Authoring core privilege expansions, collation de-escaping, and indexing constraint fixes.
  • Robert Pang: Authoring promotion timing fixes and hardening pgfincore memory safety.
  • Aniket Jha: Hardening plpgsql_check shared memory lock mechanics.
  • Sunaina Punyani: Resolving memory and execution safety in pgtt.

We also extend our sincere gratitude to the wider PostgreSQL open source communities—especially the committers, reviewers, and extension maintainers—for their collaborative reviews and shared commitment to keeping PostgreSQL the world’s most advanced open source database.

Reconnecting with the heart of open source: Highlights from our 2026 GSoC India tour

Thursday, September 17, 2026

For over twenty years, Google Summer of Code (GSoC) has welcomed new developers into open source by pairing them with experienced mentors on real projects. This spirit is especially vibrant in India, which is home to more than 55% of all global GSoC participants over the last decade.

This July, our team traveled across Bengaluru and Delhi to host a series of developer events and debut our first-ever GSoC Alumni CAMP, bringing together members of India’s vibrant GSoC alumni community. We engaged directly with current and former GSoC Contributors, Mentors, and project maintainers, experiencing firsthand the passion and energy of the Indian developer ecosystem.

Community stories: Learning to think and lead as an engineer

In both Bengaluru and Delhi, a highlight of the trip was hearing directly how open source and GSoC have fundamentally changed how developers think and work.

For many attendees, having a dedicated open source mentor through GSoC took the fear out of tackling new and intimidating codebases. One former participant told us they almost walked away from a distributed storage project because it felt too overwhelming: "Storage systems felt impossible. But great mentors taught me how to think, not just how to code." Another echoed that shift in perspective: "Why am I spending so much time thinking rather than coding? Then I realized that building products is actually about thinking more than coding. GSoC taught me to think like an engineer."

That shift in mindset turns first-time contributors into long-term open source community leaders. We met one developer who submitted their very first pull request in 2023, started mentoring in 2024, and is now a lead maintainer for a major open source Android app. We also saw how new contributors to global projects can spark entire local ecosystems—like the Indian compiler community, which started with a few GSoC alumni and has rapidly grown into a 3,500+ member network with dozens of meetups across the country.

Group photo of over a hundred Google Summer of Code alumni, mentors, and organizers wearing blue GSoC t-shirts gathered in front of a stage banner reading Google Summer of Code Alumni CAMP India 2026 in Bengaluru
The Google Summer of Code Alumni CAMP — Bengaluru

Open source mentorship in the age of AI

Across our sessions and unconference discussions, one recurring conversation resonated above all others: the evolving role of mentorship in an AI-assisted world. The human element of open source is more critical than ever. As one attendee noted:

AI can generate the slides, but the context takes nine years.

We heard over and over from participants—open source maintainer time and attention remains a limited resource. CAMP participants presented multiple examples of where AI is proving effective for generating starter templates, writing tests, or fixing syntax. Even with this, what open source projects fundamentally need hasn't changed: maintainer time, clear architectural vision, and thoughtful code reviews.

Beyond code quality, the industry agrees that dedicated mentorship is the vital bridge between temporary contributions and long-term project stewardship. Without structured guidance, newcomers often struggle with unwritten project norms, complex codebase histories, or public review feedback, leading to contributor burnout and abandoned pull requests. Programs like GSoC transform casual interest into a sustainable maintainer pipeline by fostering psychological safety, belonging, and accountable relationships. By investing directly in maintainer time and human connection, GSoC ensures that open source projects remain secure and resilient for generations to come.

What’s next?

If our trip across Bengaluru and Delhi taught us anything, it’s that the strength of open source has always come from the communities we build together, not the volume of code any one person can ship.

As developer tools evolve, our main focus for GSoC is preserving the mentorship experience that makes the program special. Manually sifting through low-quality, automated submissions wastes maintainer time and drains the energy of volunteers who signed up to mentor new peers and colleagues. We're ready to tackle these challenges directly by optimizing our program to assist and protect our community of open source maintainers, so they can focus on leading their open source projects and helping new engineers grow. You can stay updated on GSoC’s program rules and timelines at g.co/gsoc.

To everyone who joined us in Bengaluru and Delhi—thank you for your energy, your endless inspiration, and your dedication to open source!

How much should you trust your OSS data?

Thursday, September 3, 2026

 by Sophia Vargas, Google Open Source & Andrew Nesbitt, Ecosyste.ms

Every second, open source contribution quietly shapes the software we rely on, and yet our view of this open ecosystem is surprisingly opaque. Open source development is performed in public spaces — we can see the commits, issues and comments, the APIs and endpoints are free to use — the logs are just sitting there, so why can’t we just collect all of the data?


…Said every researcher, everywhere. However in most cases of open source related data, we are only looking at part of the whole. Why am I writing this post? Because many of us (including many business decision-makers) are too comfortable with unsubstantiated data. We’ve gotten used to it. Our models assume that it's smelly and we adjust the logic and weights to compromise. When it comes to open source, our confidence is even lower, even though our resulting decisions can directly impact individuals whom we collectively depend on.


Let’s consider one of my favorite datasets: GHarchive. Started as a hobby project in 2011, this crawler has amassed more than 15 years of event data from GitHub. While this source provides a historical record of open source development on GitHub, as a real-time or comprehensive source of metrics, it's unreliable and should not be a source for volume-based metrics. 


In 2025, GHarchive captured 14% fewer events than in 2024, despite steady growth in platform adoption.  Since 2025, we estimate that data retention in GHarchive has fallen to ~50% and in 2026 it may be as low as 20% for some event types (see figure below). Prior to 2025, you could make the general assumption that the majority of events would be represented in this pipeline. Since 2025, we must now assume we may be missing at least half of events and possibly more — not to mention all of the additional activity that’s left out of the event API (see GitHub’s GraphQL API.) 

The crawler logic behind this dataset is simple: give me all the events from the GitHub Event stream (e.g. opening pull requests, commenting on issues etc). However, the GitHub API has limitations on the number of calls per hour as well as the number of events listed, so for days with a lot of spiky activity, the crawler will miss some. Although we never assumed that this dataset was collecting 100% of events, the current architecture is showing signs of strain. We suspect that this is due, in part, to the rate of repository growth and adoption of automated tooling on GitHub. In 2011, GitHub announced it reached 2 million public repositories, and by 2026, that figure surpassed 400 million.  


I want to acknowledge that building and sharing comprehensive open datasets at scale is hard. Have you ever built a pipeline only to discover that the variables changed mid year, the payload for one output is getting truncated, all your joins broke because one side of the dataset is case sensitive … I could go on. And these examples are just ordinary data issues. Building a dataset at the scale of GitHub where “Every second, more than one new developer on average joined GitHub—over 36 million in the past year”—you start running into a new set of challenges.


My own journey with open source related data began when I repeatedly found myself questioning how much we could trust our own metrics. To expand my understanding of the nuances and the limitations of open source related datasets, I reached out to Andrew Nesbitt, who has spent years digging in data trenches for the benefit of the community. Together we converged on the following issues that we wanted to highlight for the broader community.


Assembling: Assume there will be problems

When I asked Andrew ‘can you summarize the challenges you have faced assembling comprehensive datasets?’ —“I just assume I'm going to have a terrible time anyway, so I start with my best effort and fill in the gaps”. While disappointing, this aligned with most data aggregation methods I’ve reviewed—tools such as Grimoire labs and OSS insights also require multiple processes for collection, combination and reconciliation. Even with these approaches, many sources have missing, incomplete, or inconsistent information.


One source is probably not enough. If you are considering the use of an open source project, you may want to know how many maintainers work on this project, what versions are available, what their dependencies are and any active vulnerabilities or known issues. Each of these queries requires a distinct source—the development history, the dependency graph, the CVE database, etc. Ecosyste.ms strives to pull this information together into one place, but combining data from 1000+ datasets has its own unique set of challenges.


For example, my index is probably not your index. One perennial issue is inconsistent naming conventions across sources. Beyond variable type and format, repository names, versions, packages, tags, licenses, urls, etc. tend to be unique across platforms. Some are case sensitive, there are often duplicates, and anyone can change a name at any time… I’ve been keenly following the adoption of purl and SWHID, but so far I have not found one name to rule them all.


Now we have to keep this up to date: At the moment, there is no consistent way of sharing updates across platforms. Changes to names, APIs, deletions, etc. are more often discovered by errors and breakage than by scouring release notes. To keep Ecosyste.ms up to date, Andrew has written multiple syncing processes that identify or infer updates that need to be accounted for. I asked Andrew ‘If you could ask a platform/data source to change one thing, what would it be?’, “Can I crawl an endpoint that's just NEW stuff?’


Consuming: Design your pipeline for your use case

Because of LLMs, “it's now easier for anyone to try to access and build reports”. But those building quick reports are likely not going to go through the pain of being comprehensive. This is where aggregated sources like GHarchive and Ecosyste.ms thrive. As data providers, we’d love if data consumers knew that:


How you collect data matters. If everyone wanted the same dataset, in the same format, at the same time, it would be simple. Depending on how the data is stored—centralized vs distributed and cached, relational vs graph, etc. —queries could be more efficient (in cost and computation) than exports or bulk requests faster than individual requests. This all depends on the topology of the infrastructure and the dataset. In a perfect world, data producers would design their architecture for their top user journeys. However open source related datasets serve a wide variety of user personas from corporations to non-profits, researchers to individual users, maintainers, funders, and many more, with a variety of demands from historical deep dives to realtime feedback. Data producers can’t design for all of these cases, so my challenge to them is to be more open about the best way to access this information. 


At the end of the day, we have to respect the human infrastructure: Open source-related datasets are riddled with personally identifiable information (PII). Some individuals may be comfortable sharing their information with fellow contributors, but seeing it aggregated across platforms can be uncomfortable. Any source with PII should be handled with care: anonymize when you can and ensure you are in alignment with policies and regulations. Open source communities are real people so please, consume their data responsibly.


Interpreting: Never stop asking questions

While many have moved on from ‘data-driven’ to ‘AI-enabled’, the fact remains that ALL AI SYSTEMS DEPEND ON DATA. Our data about open source will continue to be incomplete and imperfect, but by asking questions about our sources, acknowledging the gaps, and considering both the technical and human processes behind open source development, we can refine and improve on how we interpret our insights and models even if they don’t completely reflect reality.


Securing the agentic era: Introducing formal verification for CEL

Tuesday, August 18, 2026

CEL Formal Verification header graphic

We are rapidly entering an era where AI agents can autonomously draft, refactor, and deploy policies that protect our users and our systems. But this velocity introduces a vital question: How do we trust AI-generated policies?

Unit tests may fail to cover the infinite set of possible inputs that occur in production; thus, an AI agent that overfits its policy to existing tests may fail spectacularly in production. To secure automated policy authoring, we must combine heuristic testing with mathematical proofs.

We are thrilled to announce the Common Expression Language (CEL) Formal Verification Framework is now available. Powered by the Z3 theorem prover, this framework allows you to prove the correctness of your CEL expressions and policies, serving as the ultimate safety net for the agentic policy.

Automated reasoning definitively answers questions like:

  • “Is there any combination of inputs that allows an unapproved request into production?”
  • “Are we absolutely certain this AI-refactored policy matches the original behavior?”
  • “Can a bad actor manipulate this rule to force an evaluation error?”

Formal verification establishes mathematical certainty across the infinite spectrum of inputs. Proven policies protect your users and system while giving auditors clear proof of compliance.

To see these capabilities in action, watch our video demonstrating how the CEL Verifier REPL catches subtle logic flaws in seconds:

Proving rules from the ground up

Getting started with formal verification doesn’t require learning complex architectures right away. You can evaluate simple standalone CEL expressions to catch edge cases that tests easily miss.

(Note: The examples below use our interactive REPL syntax—check out the REPL documentation to follow along!)

1. Catching logic bugs in simple expressions (Equivalence)

How do you guarantee a refactored rule behaves identically to the original? Suppose we have a policy that allows ports 80 or 443 in production. An agent might factor the is_prod check like so:

equiv
  (is_prod && port == 80) || (is_prod && port == 443) 
  <=>
  is_prod && port == 80 || port == 443

Because logical AND has a higher operator precedence than OR, the verifier immediately flags Violated, and outputs the exact exploit: in a non-production environment (is_prod = false), the rule mistakenly allows port 443. Fixing the grouping parentheses returns Verified.

2. Enforcing exhaustive guardrails (Validity)

This capability scales directly to use cases like Kubernetes Validating Admission Policies. Suppose an engineer writes a guardrail expression that assumes every request will either be on a low port (under 80) or a high port (over 1024):

valid request.port > 1024 || request.port <= 80

When we check validity (whether an expression holds true for all inputs), the verifier exhaustively searches the entire integer space, flags Violated, and outputs the exact counterexample:

[VIOLATED] Condition is not always true. Counterexample input:
  request.port = 81

3. Guaranteeing security invariants with CEL Policy

While the verifier works perfectly with standalone CEL expressions, complex environments compose multiple rules and variables. Here, the CEL policy format shines. Using assume and assert blocks, the verifier proves a mathematical implication: if the assumptions hold, the assertions must also hold.

name: workload_admission
rule:
  variables:
    - is_admin: 'request.auth.claims.groups.exists(g, g == "admin")'
  match:
    # A subtle flaw introduced during authoring:
    - condition: 'request.is_privileged && request.is_prod'
      output: 'true'
    - condition: 'variables.is_admin || request.has_approval'
      output: 'true'
    - output: 'false'
verification:
  invariants:
    - id: universal_no_unapproved_privileged_prod
      assume:
        - 'request.has_approval == false'
        - 'variables.is_admin == false'
      assert:
        - 'rule.result == false'

The first condition admits privileged workloads into production without checking for approval or admin status. The verifier flags this and provides an example that exploits the issue:

Invariant 'universal_no_unapproved_privileged_prod' violation detected. Counterexample input:
  request.is_privileged = true
  request.is_prod = true
  request.has_approval = false
  request.auth.claims.groups = []

Assertions and assumptions define the boundaries of acceptable agent behavior, allowing developers to configure CI/CD pipelines to validate AI-generated changes simply and securely.

Under the hood: High-fidelity mathematical modeling

Translating a dynamic language into the Satisfiability Modulo Theories (SMT) domain requires immense engineering rigor to prevent the solver from hanging or hallucinating bugs. Our engine provides:

Zero false positives via three-pass taint tracking

Traditional verification tools are prone to “solver hallucinations”—reporting fake bugs when encountering custom domain-specific functions or external variables they don’t fully understand. To eliminate this noise, if a potential issue relies on an unmapped custom function, the verifier isolates and flags it as Inconclusive rather than breaking your CI pipeline with a false alarm. This guarantees every Violation report is a 100% real, reproducible bug.

Deep structural extensionality

The Formal Verification Framework offers configurable-depth bounded-model checking to prevent infinite loops within SMT quantifiers. These configurable limits allow you to control the cost of verification when analyzing deep structure equivalence in expressions like [[1], [2]] == [[1], [2]].

The mandatory bridge of trust

In the agentic era, code writes code. Mathematical proof isn’t just a nice-to-have; it is the fundamental bridge of trust developers require to let AI operate autonomously in their most sensitive systems. Get started with the CEL Formal Verification Framework, to take the next step toward a more secure agentic future today!

Let us know what you think—issues, pull requests, and feedback are always welcome!

Google joins the OpenROAD Initiative as principal member to accelerate open source silicon innovation

Tuesday, August 11, 2026

Google is committed to advancing open source silicon innovation. We are excited to share that we have formally joined the OpenROAD Initiative (ORI), Inc. as a principal member. ORI is a nonprofit public benefit corporation dedicated to the open source electronic design automation (EDA) ecosystem. As part of this commitment, Aaron Cunningham has been appointed to the ORI Governing Board to represent Google and help drive the foundation’s strategic direction, financial sustainability, and technical stewardship.

Driving long-term open source sustainability

The OpenROAD Initiative’s mission is to advance and sustain the open source EDA ecosystem by fostering collaborative innovation across research, education, and industry—transforming ideas into silicon. Google’s membership aligns directly with ORI’s multi-year sustainability goals, supported by the US National Science Foundation’s (NSF) Pathways to Enable Open-Source Ecosystems (POSE) program.

With Google’s participation and membership commitment, ORI will continue to strengthen, grow, and sustain its open source ecosystem through key vectors:

  • Neutral Stewardship: Fostering transparent governance where no single company has outsized control over the code, ensuring the project remains inspectable, accessible, and community-driven.
  • Ecosystem Growth: Supporting open and reproducible silicon research, developing robust design flows, and hosting global design contests.
  • Workforce Development: Supporting global silicon skilling initiatives by expanding open source chip design curricula and collaborating with academic institutions and industrial training networks.
  • Technical Strengthening: Enhancing continuous integration and deployment (CI/CD) pipelines, expanding PDK enablement, and improving user experience.

Leadership perspectives

“The OpenROAD Initiative is built on the vision of making chip design open and accessible to all—building a collaborative ecosystem driven by transparency and shared innovation,” said Andrew Kahng, board member of the OpenROAD Initiative. “Google’s deep commitment to open source software and hardware makes them an ideal partner. By joining at our highest membership tier, Google is helping to ensure that the open source EDA ecosystem has the stable, long-term governance and financial foundation required to grow.”

“Cutting-edge silicon research requires robust, inspectable, and reproducible toolchains,” said Drew Wingard, Director of Silicon Infrastructure, Tools and Methodology at Google. “OpenROAD has already made an incredible impact across academia and the broader industry, enabling many successful tapeouts. Google is proud to support the OpenROAD Initiative’s mission to scale this open infrastructure for the next generation of developers.”

About the OpenROAD Initiative and OpenROAD project

The OpenROAD Initiative, Inc. is a California-based 501(c)(3) nonprofit organization that provides governance, stewardship, and coordination for the OpenROAD ecosystem. The OpenROAD Project is an open source, autonomous digital chip design toolchain that democratizes semiconductor design, enabling a complete RTL-to-GDSII flow in less than 24 hours with no human in the loop. Grounded in academic research and referenced in over 500 peer-reviewed publications, OpenROAD has lowered the barriers to hardware innovation, enabling thousands of students, researchers, and startups worldwide to design and manufacture chips.

To learn more about the OpenROAD Project and install the toolchain, visit the new OpenROAD website.

For more information about membership tiers and the foundation’s governance, visit the OpenROAD Initiative website at www.openroadinitiative.org or contact membership@openroadinitiative.org.

.