opensource.google.com

Menu

This Week in Open Source for October 2, 2026

Friday, October 2, 2026

This Week in Open Source banner graphic

This Week in Open Source for October 2, 2026

A look around the world of open source

As AI becomes part of everyday software development, the most compelling questions in open source are shifting from the models themselves to everything surrounding them. I’m always drawn to how technology affects individuals, and especially the workers who use and maintain it, for better or for worse (I definitely prefer better). That’s why I’m convinced the next phase of open-source AI won’t be won by chasing bigger benchmarks, but by building the open infrastructure, transparent local tooling, and community norms that keep developers and maintainers in control. This week’s reads explore what that looks like in practice.

Upcoming Events

🗓️ October 2026

  • ValkeyConf 2026 (October 5, 2026) — Prague, Czechia. Dedicated community conference advancing the open source Valkey in-memory data store, featuring a keynote by Valkey TSC member and Google Cloud engineer Jacob Murphy on open source in-memory data structures and scaling high-performance workloads.
  • MCP Dev Summit Toronto (October 5–6, 2026) — Toronto, Ontario, Canada. Dedicated Linux Foundation developer summit advancing the Model Context Protocol (MCP) and open agentic interoperability standards. Join Google OSPO's Daryl Ducharme (that's me!) on Tuesday, October 6 (11:00 AM EDT) for "Tag-Team Transmission: Navigating A2A and MCP for Optimum Orchestration," examining how the Agent2Agent (A2A) protocol and MCP interoperate across multi-agent architectures.
  • Open Source Summit Europe (October 7–9, 2026) — Prague, Czechia. The premier Linux Foundation gathering in Europe celebrating the 35th anniversary of Linux and connecting developers, technologists, and community leaders—including a keynote by Google OSPO's Erin McKean on Docsy and technical documentation for humans and AI agents.
  • Community Over Code (October 11–14, 2026) — Glasgow, Scotland. The flagship Apache Software Foundation conference bringing together project maintainers, committers, and users to collaborate on open governance, data architecture, and community-led software development.
  • BazelCon 2026 (October 13–15, 2026) — Amsterdam, Netherlands. Annual gathering of the Bazel community uniting build system engineers, maintainers, and ecosystem contributors around fast, reproducible, multi-language software builds.
  • All Things Open 2026 (October 19–20, 2026) — Raleigh, North Carolina, USA. One of the largest community-focused open source conferences on the U.S. East Coast exploring open source software, AI engineering, and maintainer sustainability—featuring a Google keynote on Generative UI and the Open Future alongside the Google Community Lounge with the Flutter team.
  • PyTorch Conference North America 2026 (October 20–21, 2026) — San Jose, California, USA. Two days of open source machine learning innovation covering training, inference, compiler toolchains, and hardware heterogeneity across the PyTorch ecosystem.
  • AGNTCon + MCPCon North America 2026 (October 22–23, 2026) — San Jose, California, USA. Linux Foundation conference focused on building and scaling production agentic AI systems with open standards, observability, and control, featuring a keynote by Google's Rao Surapaneni.
  • GitHub Universe 2026 (October 28–29, 2026) — San Francisco, California, USA & Virtual. Annual global developer gathering highlighting open source workflows, collaborative security, and AI-assisted software engineering.

🗓️ November 2026

  • Open Source in Finance Forum (OSFF) New York (November 4–5, 2026) — New York, New York, USA. Dedicated industry conference examining open source compliance, supply-chain security, and collaborative innovation across regulated financial institutions.
  • KubeCon + CloudNativeCon North America 2026 (November 9–12, 2026) — Salt Lake City, Utah, USA. The Cloud Native Computing Foundation's flagship conference uniting adopters and maintainers around Kubernetes, platform engineering, and cloud-native infrastructure.
  • SFSCON (South Tyrol Free Software Conference) 2026 (November 13–14, 2026) — Bolzano, Italy. One of Europe's longest-running Free Software conferences bringing together public-sector decision-makers and developers to advance digital sovereignty and open infrastructure.

Open Source Reads and Links

Which of these stories will you be chatting about at your next meetup or conference? Let us know! Share with us on our @GoogleOSS X account or our @opensource.google Bluesky account.

Standardizing fine-grained access control in Apache Iceberg's REST catalog

Thursday, October 1, 2026

The Apache Iceberg community recently merged a change to the REST Catalog specification that gives lakehouses something they have long been missing: a standard, engine-neutral way to express column masking and row filtering. It is a small change, but it addresses a structural gap in how open table formats handle data governance. This post walks through the problem, the design, and why the details matter.

Finally a single fine-grained access control standard for any engine

The problem: there is no server in the read path

In a traditional database, access control is straightforward because there is exactly one door. Every query passes through the database server, the server knows who is asking, and if policy says "this user only sees the last four digits of the card number," the server applies that policy before returning results. One process, one enforcement point.

Apache Iceberg deliberately removed that door. The data is Parquet files in object storage, and any engine—such as BigQuery, Spark, Trino, Flink, PyIceberg, or DuckDB—can read those files directly. This design is what makes Iceberg fast and interoperable. But it also means that once an engine asks the catalog "where is the payments table?" and receives the metadata pointer, it has full access to the files. The catalog can grant or deny access to the whole table, but it has had no vocabulary for "access granted, but mask the email column" or "access granted, but only on rows where region = 'US'."

Until now, Iceberg had no built-in answer to this. The REST spec did offer coarser instruments like credential vending, which controls storage access per table, and server-side scan planning, which lets a catalog withhold entire data files. But neither can express "mask this column" or filter rows that are not already physically separated into their own files. So users who needed fine-grained access control had to step outside the open protocol entirely. They could adopt a vendor-provided client that understands its proprietary policy format, or route every read through a vendor's proxy service, giving up the direct-to-storage performance that motivated the use of Iceberg in the first place. It also quietly undermines Iceberg's core promise: the moment governance requires a specific vendor's client, the table is no longer open to any engine.

The new read-restrictions field in the REST Catalog spec is the community standardizing that vocabulary.

The mechanism

When an engine calls loadTable, the response may now include an optional read-restrictions object:

Read restrictions payload showing required column projections and required row projections

Read restrictions are expressed in two fields:

  • required-row-filter: a standard Iceberg predicate expression. Rows for which it evaluates to false must not appear in the result, and no information derived from them may be included.
  • required-column-projections: a list of columns, identified by field ID, each with a transformation the reader must apply before returning values.

One evaluation rule ties them together: the row filter is evaluated against the original, untransformed column values, and projections are applied to the rows that survive. This ordering is what makes the two features composable. A policy can filter on region = 'US' and also mask region in the output, and the filter still works. If masking ran first, any policy that filtered on a masked column would silently break.

Note what the catalog is doing here. It evaluates the access policy server-side—it knows the caller's identity from the authentication token—and returns only the result of that evaluation. The policy itself, with its roles, tags, and governance model, never crosses the wire. The engine does not need to understand how any particular catalog models governance. It needs to understand nine actions and a predicate.

The whole architecture fits in one sequence:

Sequence diagram illustrating the Lakehouse credential vending and reader-side fine-grained access control workflow between an End User, a Trusted Engine (BigQuery/Spark), the Lakehouse runtime catalog, and Google Cloud Storage (GCS).
The catalog stays the policy decision point; the trusted engine becomes the policy enforcement point; the end user never holds storage credentials. Sections below unpack the three load-bearing details in this picture: the enforcement ordering, the fail-closed branch, and the trust boundary.

The nine masking actions

The specification defines a closed set of nine masking actions, each with an exact, per-type definition. The goal is cross-engine consistency: Spark, Trino, and PyIceberg must produce identical output for any given masking action. These actions are being implemented in iceberg-core, providing engines with spec-compliant transformations out of the box rather than requiring them to reimplement byte-level logic independently.

Below, the actions are grouped by the analytical utility of the resulting masked data:

Preserve the shape, hide the value

  • mask-alphanum—digits become n, other characters become x, with a small allowlist of punctuation (( ) , . - @) preserved. iceberg16112018@apache.org becomes xxxxxxxnnnnnnnn@xxxxxx.xxx—recognizably an email address, but not whose.
  • show-first-4 / show-last-4—preserve four code points, and apply mask-alphanum to the rest. 4111-1111-1111-4444 becomes nnnn-nnnn-nnnn-4444, the familiar customer-support view of a card number.

Hide everything

  • replace-with-null—the value becomes NULL. Only valid for optional fields; a server must not return it for a required field, and a reader that receives one must fail the query.
  • mask-to-fixed-value—the value becomes a type-specific constant (0, "XXXXXXXX", the epoch, an all-zero UUID, an empty list, and so on, each spelled out in the spec). Uniquely among the actions, this one also replaces NULL inputs—even the null-or-not bit is hidden.

Reduce precision

  • truncate-to-year / truncate-to-month—2024-07-15 becomes 2024-01-01 or 2024-07-01. Cohort analytics keeps working; identifying individuals gets harder.

Allow joins, hide values

  • sha-256-global—deterministic SHA-256, with exact byte-encoding rules per input type. The same input always produces the same output, everywhere, so GROUP BY user_id and joins across tables on a hashed key still work. The cost of that determinism is that hashed values can be tested against precomputed guesses—this is pseudonymization, not encryption.
  • sha-256-query-local—the same hash, salted with a fresh, cryptographically random salt (at least 16 bytes) per query. Values remain consistent within a single query, so self-joins and aggregation work, but cannot be correlated across queries, and precomputed-guess attacks no longer apply.

This last pair is a nice piece of design: the tradeoff between linkability and privacy, expressed as two enum values that a policy author chooses between per column.

Every action produces a value of the same type as its input—masked strings are strings, truncated dates are dates—so restrictions never change the schema an engine plans against. Queries do not need rewriting; values simply arrive transformed.

Fail-closed by design

The most consequential sentence in the specification is this one:

If a trusted reader that supports read-restrictions cannot apply any returned restriction, it must fail the query and must not silently return raw, partial, or empty results.

Consider the failure modes. A catalog sends an action added in a future spec version that the engine does not recognize. Or an expression type it cannot evaluate. The convenient behavior would be to skip what it does not understand and return the data. The spec rules this out: unrecognized action, fail; unparseable filter, fail; duplicate field ID in the projections, fail. Every ambiguity resolves to "no data" rather than "raw data." This is the right default for an access control mechanism, and it is also the one implementers would be tempted to soften—which is exactly why it is normative in the spec rather than left to judgment. It is also what allows the action vocabulary to grow in future versions without older engines becoming silent leak vectors.

A few prohibitions in the spec reward a closer look, because each one closes a subtle correctness hole:

  • No projections on map keys. Masking a map's keys can collapse two keys into one, or produce null keys, which engines silently coalesce or reject—data corruption presented as privacy. The spec bans it outright.
  • No projection on both a nested type and a field inside it. Masking a struct and also a field within that struct has no well-defined order of operations, so the spec refuses to define one: servers must not send it, and readers must reject it.
  • Everything references field IDs, never column names. This is standard Iceberg discipline: if ssn is renamed to national_id, the policy remains bound to the same physical column. A name-based policy would silently detach on rename—the worst possible failure mode for access control.

The trust model, stated plainly

All of this is enforced by the reader. The catalog hands the engine the file locations along with the restrictions, and a client that chooses to ignore the restrictions can read the raw files. So what is this actually protecting?

The specification is explicit: this mechanism assumes a trust relationship between the catalog and the engine, and how that trust is established is deliberately out of scope. The intended deployment is one where a platform team's engines—the shared Spark and Trino clusters—are trusted enforcement points that hold storage access (for example, through credential vending), while end users only ever talk to those engines and never hold storage credentials themselves. The trusted engine becomes the enforcement point, playing the role the database server played in the traditional architecture. The difference is that its behavior is now defined by a common, open protocol rather than by N proprietary integrations.

In other words, read restrictions do not protect data from the engine; they let the catalog direct a trusted engine to protect data from the engine's users. For genuinely untrusted readers, the coarse-grained model still applies: they are restricted to vended credentials where the storage access granted to the user aligns with the data permission of the user.

One operational subtlety deserves attention: restrictions are per-response and per-identity. The same loadTable call made by a different principal—or by the same principal later—may return different restrictions, and the restrictions attach to every read performed with that response, including subsequent planTableScan and fetchScanTasks calls. The spec therefore requires that the response not be cached outside its authentication scope. If your platform caches loadTableResponse, that cache is now security-sensitive and worth an audit.

What is still missing

Read restrictions are a foundation, not the finished building. It is worth being honest about the gaps between this specification and complete fine-grained access control:

  1. The action vocabulary is fixed, because Iceberg Expressions are not implemented yet. Nine actions cover the common masking policies, but they are a deliberately closed set: a policy like "apply this custom redaction function" cannot be expressed, and the row filter is limited to predicates—comparisons that produce true or false. The path to generalizing this already exists on paper: the Iceberg Expressions specification, proposed by Ryan Blue and adopted in mid-2026, defines a portable structure for value expressions—constants, field references, and calls to well-defined functions or SQL UDFs. Once engines can evaluate those expressions, a catalog could return arbitrary transformations instead of choosing from an enum. Today no engine implements general expression evaluation, which is why the initial design confines itself to a small vocabulary that can be specified bit-for-bit.
  2. There is no policy definition—deliberately. The specification standardizes the result of policy evaluation, never the policy itself. How an organization expresses "analysts see masked PII, auditors see everything"—the roles, tags, rules, and administrative APIs—remains entirely the catalog vendor's domain, and the assumption is that it stays there. Only the consequences of a policy are portable across engines; the policy definition is not. Whether communities eventually want a portable policy format is an open question the spec does not attempt to answer.
  3. Trusted clients are asserted, not proven. As the trust-model section noted, how a catalog establishes that a caller is a trusted, enforcing engine is out of scope. In practice that trust is deployment configuration—service identities, network boundaries, and which principals receive vended credentials. There is no attestation mechanism in the protocol by which an engine proves it enforces restrictions; the trusted-client mechanism is an assumption the platform operator must make true.

None of these gaps undermines the design—each is a deliberate scoping decision that kept the proposal small enough to reach consensus—but they define the roadmap for what "complete" fine-grained access control in the open lakehouse still requires.

Why this matters

The specification change itself does not ship enforcement; that work in the engines begins now, starting with the default actions. But the shape of the design is right in three ways:

  1. It picks the honest enforcement point. In an architecture with no server in the read path, the trusted engine is the only place enforcement can live without giving up direct storage reads. The spec accepts that constraint explicitly rather than obscuring it.
  2. It standardizes the narrow waist. Catalogs keep their own rich policy engines—roles, tags, and attribute-based rules. Engines implement nine actions and a predicate evaluator, once. N×M becomes N+M.
  3. It is fail-closed everywhere, which is what makes the vocabulary safely extensible.

The pattern—the catalog evaluates policy against the caller's identity and returns a small, closed vocabulary of obligations that the client must enforce or fail—is a useful template, and it would not be surprising to see more of the governance surface expressed this way over time.

The proposal was approved on the Apache Iceberg dev list with eight binding +1 votes and no objections. This is a strong signal of consensus across the many companies and open source communities that participate in the project. Consistent, engine-independent enforcement of fine-grained policies is a property the ecosystem has long wanted; it now has a specification for it, and the interesting work of implementing it in engines and catalogs is underway. If you work on either, the dev list is the place to get involved.

The full schema is available in rest-catalog-open-api.yaml under ReadRestrictions. View the full vote thread.

New public datasets available in Google Cloud Lakehouse

Wednesday, September 30, 2026

An image of a castle with an ice block coming out of the front gate

Did you know the Google Chrome Wikipedia page received over 36 million page views in 2025, or that Chromium logged over 3,700 commits in April 2010? If you have needed large, real-world datasets to benchmark query engines on Apache Iceberg, Google Cloud’s Lakehouse team is excited to announce the release of new public datasets in Google Cloud Lakehouse to help you explore and analyze open data at scale.

Google Cloud’s Lakehouse provides a high-performance storage catalog using Apache Iceberg as its open table format. By decoupling storage from compute, it enables you to use your preferred query engines—such as BigQuery, Apache Spark, or Trino—while managing data in an open format to avoid vendor lock-in.

These new public datasets are designed to help you explore the Apache Iceberg ecosystem and begin working immediately with real-world data. We are providing access to some of the most popular BigQuery public datasets, including Wikipedia pageviews and GitHub commit histories. Let’s dive into how you can start querying them.

Prerequisites

  • A Google Cloud project (for authentication).
  • Standard Google Application Default Credentials (ADC) configured in your environment.

Explore metadata with PyIceberg

To explore table metadata using PyIceberg, install the required Python packages in a virtual environment:

sudo apt-get install python3-venv
python3 -m venv public_data_lakehouse
source public_data_lakehouse/bin/activate
pip install google-auth
pip install pyiceberg
pip install pyarrow

After installing the required packages, you can inspect a table’s schema using the Python script below. Here is how to inspect the table containing Wikipedia page view data from 2016:

from pyiceberg import catalog as pyiceberg_catalog

PROJECT = "<YOUR_PROJECT_ID>"
NAMESPACE = "wikipedia"
TABLE = "pageviews_2016"

def load_catalog() -> pyiceberg_catalog.Catalog:
  properties = {
      "type": "rest",
      "uri": "https://biglake.googleapis.com/iceberg/v1/restcatalog",
      "warehouse": "gs://lakehouse-public-data",
      "auth": {"type": "google"},
      "header.X-Iceberg-Access-Delegation": "vended-credentials",
      "header.x-goog-user-project": PROJECT,
  }
  return pyiceberg_catalog.load_catalog(PROJECT, **properties)

def main() -> None:
  catalog = load_catalog()
  table = catalog.load_table((NAMESPACE, TABLE))
  print(f"Schema for `{NAMESPACE}.{TABLE}`:", table.schema())

if __name__ == "__main__":
  main()

Note: Replace <YOUR_PROJECT_ID> with your actual Google Cloud Project ID. This is required for the REST catalog to authenticate your quota usage, even for free public access.

PySpark and Managed Service for Apache Spark

Now that we have explored the table’s schema, we can use a managed, serverless Spark notebook to query the tables. This provides the speed and flexibility of Apache Spark without creating or managing a cluster. You can run this script in Google Cloud's serverless Managed Service for Apache Spark (formerly Dataproc):

from google.cloud.dataproc_v1 import Session
from google.cloud.dataproc_spark_connect import DataprocSparkSession

PROJECT_ID = "<YOUR_PROJECT_ID>"
spark_catalog = "lakehouse-public-data"

session = Session()
session.runtime_config.properties = {
  "spark.sql.extensions": "org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions",
  f"spark.sql.catalog.{spark_catalog}": "org.apache.iceberg.spark.SparkCatalog",
  f"spark.sql.catalog.{spark_catalog}.type": "rest",
  f"spark.sql.catalog.{spark_catalog}.uri": "https://biglake.googleapis.com/iceberg/v1/restcatalog",
  f"spark.sql.catalog.{spark_catalog}.rest.auth.type": "org.apache.iceberg.gcp.auth.GoogleAuthManager",
  f"spark.sql.catalog.{spark_catalog}.io-impl": "org.apache.iceberg.gcp.gcs.GCSFileIO",
  f"spark.sql.catalog.{spark_catalog}.header.X-Iceberg-Access-Delegation": "vended-credentials",
  f"spark.sql.catalog.{spark_catalog}.warehouse": "gs://lakehouse-public-data",
  f"spark.sql.catalog.{spark_catalog}.header.x-goog-user-project": PROJECT_ID,
}

spark = (
   DataprocSparkSession.builder
     .appName("Lakehouse Public Data Demo")
     .dataprocSessionConfig(session)
     .getOrCreate()
)
spark.conf.set("spark.sql.defaultCatalog", spark_catalog)

With the session active, you can query monthly Wikipedia view counts for BigQuery-related articles by executing the following code in a new cell:

df1 = spark.sql("""
SELECT
  title,
  wiki,
  SUM(views) AS total_views
FROM wikipedia.pageviews_2026
WHERE datehour >= TIMESTAMP '2026-01-01 00:00:00'
  AND datehour <  TIMESTAMP '2026-02-01 00:00:00'
  AND LOWER(title) LIKE '%bigquery%'
GROUP BY title, wiki
ORDER BY total_views DESC
LIMIT 20
""")
df1.show(10)

Or, you can retrieve recent GitHub commits referencing Iceberg using the following snippet:

df2 = spark.sql("""
SELECT
  commits.commit,
  commits.subject,
  commits.message,
  commits.author.name AS author_name,
  timestamp_seconds(commits.committer.date.seconds) AS commit_time,
  repo_name
FROM github_repos.commits AS commits
WHERE LOWER(commits.subject) LIKE '%iceberg%'
   OR LOWER(commits.message) LIKE '%iceberg%'
ORDER BY commit_time DESC
LIMIT 50
""")
df2.show(10)

Start building today

These datasets were imported from their BigQuery counterparts and transformed using Apache Iceberg as the table format and Parquet as the data file format. Our goal is to lower the entry barrier so you can learn and explore Apache Iceberg with your favorite query engine without managing any infrastructure. To get started with building an open, managed, and high-performance Iceberg lakehouse, visit the Google Cloud Lakehouse page.

This Week in Open Source for September 24, 2026

Thursday, September 24, 2026

This Week in Open Source banner graphic

This Week in Open Source for September 24, 2026

A look around the world of open source

Open source underpins an $8.8 trillion global economy, yet sustaining it requires moving beyond voluntary charity and philosophical appeals. This week’s reads examine how the ecosystem is adapting to modern economic and AI pressures—from proposing package-registry royalties and quantifying 2–5x net value for regulated utility grids, to turning autonomous AI security incidents into hardened supply-chain defenses that empower open-science breakthroughs like NASA and IBM’s Lunar Foundation Model.

Upcoming Events

🗓️ October 2026

  • MCP Dev Summit Toronto (October 5–6, 2026) — Toronto, Ontario, Canada. Dedicated Linux Foundation developer summit advancing the Model Context Protocol (MCP) and open agentic interoperability standards. Join Google OSPO's Daryl Ducharme on October 6 for "Tag-Team Transmission: Navigating A2A and MCP for Optimum Orchestration," examining how the Agent2Agent (A2A) protocol and MCP interoperate across multi-agent architectures.
  • Open Source Summit Europe (October 7–9, 2026) — Prague, Czechia. The premier Linux Foundation gathering in Europe celebrating the 35th anniversary of Linux and connecting developers, technologists, and community leaders across open AI, embedded systems, and digital trust.
  • Community Over Code (October 11–14, 2026) — Glasgow, Scotland. The flagship Apache Software Foundation conference bringing together project maintainers, committers, and users to collaborate on open governance, data architecture, and community-led software development.
  • All Things Open 2026 (October 19–20, 2026) — Raleigh, North Carolina, USA. One of the largest community-focused open source conferences on the U.S. East Coast exploring open source software, AI engineering, DevSecOps, and maintainer sustainability.
  • GitHub Universe 2026 (October 28–29, 2026) — San Francisco, California, USA & Virtual. Annual global developer gathering highlighting open source workflows, collaborative security, and AI-assisted software engineering.

🗓️ November 2026

  • Open Source in Finance Forum (OSFF) New York (November 4–5, 2026) — New York, New York, USA. Dedicated industry conference examining open source compliance, supply-chain security, and collaborative innovation across regulated financial institutions.
  • KubeCon + CloudNativeCon North America 2026 (November 9–12, 2026) — Salt Lake City, Utah, USA. The Cloud Native Computing Foundation's flagship conference uniting adopters and maintainers around Kubernetes, platform engineering, and modern cloud infrastructure.
  • SFSCON (South Tyrol Free Software Conference) 2026 (November 13–14, 2026) — Bolzano, Italy. One of Europe's longest-running Free Software conferences bringing together public-sector decision-makers and developers to advance digital sovereignty and open infrastructure.

Open Source Reads and Links

  • [Blog] Nobody pays for open source. We can force them to. — Companies already spend over $1 billion a year on open source, but that money goes to supply-chain mirror and security vendors rather than the maintainers writing the code. Laurie Voss breaks down why voluntary charity and license changes always fail, arguing that commercial package registries and mirror vendors should pay automatic royalties down the dependency tree to long-tail maintainers.
  • [Article] AI Is reshaping open source software and straining the systems that sustain it — As conversations grow around moderating the speed of AI development across the industry, it is critical to watch the friction points where AI and open source overlap—particularly as an influx of AI-generated code and vulnerability discovery strains maintainers of an $8.8 trillion ecosystem. Funding remains front and center, requiring not just raw capital but continuous monitoring for effectiveness in sustaining the human governance, consensus-building, and non-coding infrastructure that AI cannot replace.
  • [Report] LF Energy Research Finds Open Source Software Can Deliver 2-5x Greater Net Value for Grid Operators — While open source software and public utilities seem like natural philosophical allies as public goods, philosophical appeals routinely fail in regulated infrastructure sectors where operators are bound by strict reliability mandates and ratepayer accountability. This new LF Energy report demonstrates how to bridge that gap: by translating collaborative "make together" governance into an auditable benefit-cost framework showing two to five times greater net value over proprietary procurement.
  • [Article] AI Is at a Turning Point — High-profile incidents of autonomous AI agents breaching Hugging Face or social-engineering open source maintainers inevitably dominate headlines and amplify purely negative narratives around AI. However, treating these security failures as concrete post-mortems—exposing how standard reinforcement learning incentivizes deceptive subgoals—is exactly what allows open source communities to establish hardened supply-chain defenses and safely advance high-impact, positive AI use cases.
  • [Article] IBM and NASA Release Open-Source AI Model to Support Lunar Exploration — Seeing open source, AI, open data, and space science converge is genuinely inspiring, especially when it demonstrates that open scientific models are only as useful as the open datasets beneath them. While NASA and JAXA planetary archives have long been public, IBM and NASA's release of the Lunar Foundation Model alongside the first harmonized, machine-learning-ready lunar dataset (unifying 30+ multi-instrument layers) proves that transforming raw open data into shared, interoperable infrastructure is what unlocks domain discovery.

Which of these stories will you be chatting about at your next meetup or conference? Let us know! Share with us on our @GoogleOSS X account or our @opensource.google Bluesky account.

Google Cloud: Investing in the future of PostgreSQL — 2026 highlights

Wednesday, September 23, 2026

Google Cloud is committed to open source, and PostgreSQL is a cornerstone of managed database offerings, including Cloud SQL and AlloyDB.

Continuing our work with the PostgreSQL communities, we've been contributing to the core engine and participating in the patch review process. Below is a summary of that technical activity between January 2026 and September 2026, highlighting our efforts to enhance the performance, stability, and resilience of the upstream project and ecosystem. By strengthening these core capabilities, we aim to drive innovation that benefits the entire global PostgreSQL ecosystem and its diverse user base.

Our technical contributions in this period have focused on enhancing core engine performance, introducing features for logical replication, fixing critical bugs, and improving upgrade resilience. We also continue to invest in the PostgreSQL ecosystem by addressing bugs in widely used extensions.

Technical contributions: January 2026 – September 2026

Our contributions this cycle span four key areas:

  1. Logical replication and conflict management: Paving the way for active-active replication, enhanced conflict logging, and catalog safety.
  2. Core engine performance and reliability: Eliminating tuple-level overhead in sequential scans and making promotion timeouts accurate.
  3. Core catalog, collation, and indexing bug fixes: Strengthening privilege consistency, deferrable index builds, collation handling, and memory safety.
  4. PostgreSQL extension and ecosystem hardening: Resolving critical crashes, lock tranche registrations, and memory vulnerabilities across popular ecosystem extensions (plpgsql_check, pgfincore, and pgtt).

1. Logical replication and conflict management

Logical replication is essential for near-zero downtime migrations, major version upgrades, and multi-region data distribution. Our recent work focuses on conflict log table infrastructure, namespace clarity, and cross-database catalog hygiene.

Conflict log table infrastructure

  • Background and challenge: A major milestone on the roadmap to active-active multi-master replication is the ability to automatically record and resolve data discrepancies across nodes. Without structured conflict logs, tracking conflicting writes requires inspecting server logs or relying on custom application handlers.
  • Solution: This patch introduces the foundational conflict log table management infrastructure as an option in CREATE SUBSCRIPTION. It establishes the catalog structures and configuration hooks necessary to direct conflict records into dedicated, queryable tables, serving as the cornerstone for upcoming automatic conflict logging into tables.
  • Note: Provided no blocking issues arise, this functionality is slated for inclusion in PG 20.
  • Contributors: Dilip Kumar (Primary Author)

Fix REASSIGN OWNED for subscriptions in other databases

  • Background and challenge: While pg_subscription is physically a shared catalog so the background launcher process can scan all databases, subscription objects are logically local to each database. Certain operations, such as REASSIGN OWNED, failed to restrict their catalog scans to the current database (MyDatabaseId), leading to accidental modifications across database boundaries.
  • Solution: Added explicit guards to pg_subscription readers to ensure non-launcher processes filter strictly by MyDatabaseId, protecting cross-database isolation and updating documentation.
  • Contributors: Dilip Kumar (Author)

Schema-qualified names in EXCEPT clause error messages

  • Background and challenge: When publishing tables with EXCEPT clauses, check_publication_add_relation() previously reported only unqualified table names when a relation could not be processed, leading to ambiguous error messages in multi-schema databases.
  • Solution: Updated error reporting paths to emit fully schema-qualified relation names, aligning with PostgreSQL's broader error messaging standards.
  • Contributors: Dilip Kumar (Author)

2. Core engine performance and administrative enhancements

Optimizing throughput and improving the predictability of administrative operations remain top priorities for database workloads.

Timeout handling in pg_promote()

  • Background and challenge: Standby promotion via pg_promote() allows users to specify a timeout interval. Due to imprecise elapsed-time tracking during wait loops, promotion operations could terminate prematurely before the configured timeout elapsed.
  • Solution: Refined the promotion loop to pre-calculate the expected end timestamp and track actual elapsed time across iterations, ensuring strict adherence to user-configured wait intervals.
  • Contributors: Robert Pang (Reporter and Author)

Sequential scan performance optimization

  • Background and challenge: Previously, CheckXidAlive validation was executed inside the inner table_scan_next routines. This incurred repetitive check overhead on every single tuple fetched during sequential scans.
  • Solution: Restructured the scan control flow to eliminate redundant per-tuple checks, yielding cleaner execution paths and measurable throughput improvements on large table scans.
  • Contributors: Dilip Kumar (Author)

3. Core engine access control, collation, and indexing bug fixes

We continue to harden PostgreSQL's core engine against catalog inconsistencies, segmentation faults, and edge-case query anomalies.

Large object access with pg_{read,write}_all_data

  • Problem: The default roles pg_read_all_data and pg_write_all_data were designed to allow maintenance utilities like pg_dump to operate without superuser privileges. However, Large Objects (LOBs) remained inaccessible under these roles without explicit object-level grants.
  • Fix: Updated permission checks to extend pg_read_all_data and pg_write_all_data coverage to Large Objects, completing superuser-free dump and maintenance workflows.
  • Contributors: Nitin Motiani (Author), Dilip Kumar (Reviewer)

Immediate property propagation in index copies (REINDEX CONCURRENTLY)

  • Problem: When building a replacement index during REINDEX CONCURRENTLY for a deferrable unique constraint, index_create_copy() defaulted constraint flags to 0, setting the immediate property to true. This caused concurrent transactions to immediately trigger constraint violations rather than deferring verification until commit time.
  • Fix: Introduced the INDEX_CREATE_DEFERRABLE flag to properly propagate an immediate property of false to transient copied indexes without violating internal constraint assertions.
  • Contributors: Nitin Motiani (Author)

LIKE matching with nondeterministic collations and backslashes

  • Problem: Following the addition of nondeterministic collation support for LIKE, literal pattern substring parsing unconditionally skipped all backslash characters. When encountering escaped backslashes (\\), the engine omitted the second backslash entirely instead of emitting a literal \.
  • Fix: Corrected pattern de-escaping logic to correctly recognize and emit escaped backslashes during evaluation.
  • Contributors: Nitin Motiani (Author)

DSM lock release and crash prevention

  • Problem: If a backend encountered a FATAL exit while holding a lock in a Dynamic Shared Memory (DSM) segment (e.g., inside dynamic shared hashtables dshash) outside of an active transaction, releasing locks during process termination could reference already detached DSM segments, triggering a segmentation fault.
  • Fix: Hardened cleanup and lock release sequences during fatal exits to safely detach memory segments without segfaulting.
  • Contributors: Dilip Kumar (Reviewer)

4. PostgreSQL ecosystem and extension hardening

Enterprise PostgreSQL architectures rely heavily on third-party extensions. Our team actively contributes bug fixes and stability improvements upstream to critical ecosystem projects.

plpgsql_check: LWLock tranche registration for PG14

  • Problem: In PG14 and earlier, loading plpgsql_check via shared_preload_libraries could fail with shared memory lock errors due to missing or outdated named LWLock tranche registrations (plpgsql_check profiler funcs stats and plpgsql_check profiler func stmts stats).
  • Fix: Aligned pre-PG15 tranche initialization with modern shmem_request_hook patterns, guaranteeing safe shared memory allocation on older server versions.
  • Contributors: Aniket Jha (Author)

pgfincore: Memory safety hardening

  • Problem: pgfincore contained two subtle memory corruption issues: an off-by-one array boundary access during buffer inspection and a dangling pointer assignment during deallocation.
  • Fix: Authored patches to enforce strict boundary checks and clean pointer resets, eliminating potential memory corruption during OS buffer cache analysis.
  • PRs: klando/pgfincore#12 and klando/pgfincore#13
  • Contributors: Robert Pang (Author)

pgtt: Use-after-free prevention on cached plans

  • Problem: When running utility commands via the extended query protocol or within cached contexts (such as PL/pgSQL and SQL functions), pgtt modified cached statement parse trees in-place using short-lived query memory. Once that memory was freed, subsequent executions of the cached plan led to Use-After-Free crashes.
  • Fix: Updated the extension to operate on an isolated, deep copy of the parse tree for cached and read-only statements, ensuring memory safety across repeated executions.
  • Contributors: Sunaina Punyani (Author)

Related reading

Community roadmap: Your feedback matters

We encourage you to utilize the comments area to propose new capabilities or refinements you wish to see in future iterations, and to identify key areas where the PostgreSQL open source communities should focus their investments.

Acknowledgments

We would like to celebrate our engineers for their ongoing dedication to open source:

  • Dilip Kumar (PostgreSQL Significant Contributor): Authoring and reviewing core replication, catalog, memory, and performance patches.
  • Nitin Motiani: Authoring core privilege expansions, collation de-escaping, and indexing constraint fixes.
  • Robert Pang: Authoring promotion timing fixes and hardening pgfincore memory safety.
  • Aniket Jha: Hardening plpgsql_check shared memory lock mechanics.
  • Sunaina Punyani: Resolving memory and execution safety in pgtt.

We also extend our sincere gratitude to the wider PostgreSQL open source communities—especially the committers, reviewers, and extension maintainers—for their collaborative reviews and shared commitment to keeping PostgreSQL the world’s most advanced open source database.

.