Python is one of the primary languages at Google, powering everything from machine learning research to developer infrastructure and data pipelines. Across hundreds of thousands of files in our monorepo, fast and reliable static type checking is essential for maintaining code safety and developer velocity.
Today, Pyrefly—an open source type checker developed by Meta—is the Python type checker at Google.
Replacing our previous type-checking infrastructure with Pyrefly provided significant speedups for engineers and agents alike, while substantially reducing build-system compute resources.
Moving away from Pytype
For more than a decade, Google relied on Pytype, an internally developed type checker that pioneered Python type checking and collaborated with the open source community to create typeshed.
However, as Python’s typing system evolved rapidly, Pytype faced fundamental challenges. Because it operated by analyzing compiled bytecode rather than source ASTs, keeping pace with modern typing PEPs became an increasing maintenance burden due to bytecode instability across Python releases. In addition, it struggled to deliver the fast turnaround times required for modern development loops.
As detailed in the Pytype update, we decided that Python 3.12 would be the final supported version for Pytype, leading us to evaluate and adopt modern open source alternatives.
Why Pyrefly?
Pyrefly incorporates years of collective lessons from earlier tools across the Python typing ecosystem. When evaluating candidates to succeed Pytype, Pyrefly quickly stood out across three key areas:
Performance
Written in Rust, Pyrefly is designed for high throughput and lazy, parallel evaluation. In our internal benchmarks across various Google projects, it proved to be an order of magnitude faster than Pytype, while scaling smoothly across large dependency graphs.
Typing spec conformance
Pyrefly achieves strong conformance—scoring roughly 97% on the official typing conformance test suite—and is actively maintained to track new Python versions and typing PEPs. This comprehensive language support makes Python version upgrades across our monorepo smoother.
In addition to supporting new syntax, Pyrefly provides strict type safety in areas where Pytype was historically permissive. A notable example is unsound unions: Pytype allowed passing a value of a union type (such as int | None or int | str) to a function expecting a specific type (such as int), as long as at least one type in the union was compatible. However, Pyrefly enforces sound union checking, catching these mismatches:
# foo.py
def process_id(x: int) -> None:
...
def get_id() -> int | None:
...
val = get_id()
# Accepted by Pytype, but rejected by Pyrefly:
process_id(val)
Actionable error diagnostics
Pyrefly provides Rust-style compiler diagnostics, displaying the offending code snippet with inline annotations that point directly to the root cause of type mismatches, sometimes with suggested fixes. E.g., for the foo.py example above, Pyrefly produces:
ERROR Argument `int | None` is not assignable to parameter `x` with type `int` in function `process_id`[bad-argument-type]--> foo.py:8:12
|8 | process_id(val)
|^^^|
The declared type does not allow `None`. Consider narrowing the value with an `is not None` check.
Performance and infrastructure impact
Switching to Pyrefly brought measurable improvements across Google's developer ecosystem and infrastructure:
Up to 98% faster incremental rebuilds: In developer edit-and-rebuild workflows (with a warm daemon during active editing cycles), Pyrefly delivered up to a 98% latency reduction across various targets. On large machine learning targets, type-checking times dropped from minutes to seconds.
>90% critical-path reduction in clean builds: In cold-cache benchmark suites across major libraries and models, Pyrefly consistently reduced the type-checking share of the build critical path by 90% to 99%, eliminating a long-standing bottleneck in our build pipelines.
>80% compute hardware savings: Pyrefly reduced Google’s daily peak compute occupancy for Python type checking by more than 80%, saving thousands of machine cores every day.
To illustrate this impact on developer workflows, the chart below compares total edit-and-rebuild latency between Pytype and Pyrefly across targets of varying sizes.
Developer feedback
Beyond aggregate metrics, Pyrefly improved the day-to-day development loop, allowing engineers and AI coding agents to catch bugs faster.
Here is what engineers across Google have shared about their experience:
"I haven't waited for a Pytype action to complete since our project switched [to Pyrefly]. It is doubly awesome for agentic coding. Type checking is faster than running the tests now." — Peter Hawkins, JAX Tech Lead
"I really like that Pyrefly gives super clear errors that point out exactly what's wrong." — Yotam Doron, Gemini Large Scale Pretraining
“Investing in Python tooling pays huge dividends for research velocity: Pyrefly keeps our experimental iterations fast and catches subtle bugs early with clear, actionable errors.” — Tom Ward, GDM Science
Looking forward
Adopting Pyrefly highlights the value of uniting behind shared open source developer tooling. We are deeply grateful to the Pyrefly team for their rapid turnaround and responsiveness on upstream issues throughout our rollout. We look forward to continuing our collaboration and contributing to the Python open source community.
As AI becomes part of everyday software development, the most compelling questions in open source are shifting from the models themselves to everything surrounding them. I’m always drawn to how technology affects individuals, and especially the workers who use and maintain it, for better or for worse (I definitely prefer better). That’s why I’m convinced the next phase of open-source AI won’t be won by chasing bigger benchmarks, but by building the open infrastructure, transparent local tooling, and community norms that keep developers and maintainers in control. This week’s reads explore what that looks like in practice.
Upcoming Events
🗓️ October 2026
ValkeyConf 2026 (October 5, 2026) — Prague, Czechia. Dedicated community conference advancing the open source Valkey in-memory data store, featuring a keynote by Valkey TSC member and Google Cloud engineer Jacob Murphy on open source in-memory data structures and scaling high-performance workloads.
MCP Dev Summit Toronto (October 5–6, 2026) — Toronto, Ontario, Canada. Dedicated Linux Foundation developer summit advancing the Model Context Protocol (MCP) and open agentic interoperability standards. Join Google OSPO's Daryl Ducharme (that's me!) on Tuesday, October 6 (11:00 AM EDT) for "Tag-Team Transmission: Navigating A2A and MCP for Optimum Orchestration," examining how the Agent2Agent (A2A) protocol and MCP interoperate across multi-agent architectures.
Open Source Summit Europe (October 7–9, 2026) — Prague, Czechia. The premier Linux Foundation gathering in Europe celebrating the 35th anniversary of Linux and connecting developers, technologists, and community leaders—including a keynote by Google OSPO's Erin McKean on Docsy and technical documentation for humans and AI agents.
Community Over Code (October 11–14, 2026) — Glasgow, Scotland. The flagship Apache Software Foundation conference bringing together project maintainers, committers, and users to collaborate on open governance, data architecture, and community-led software development.
BazelCon 2026 (October 13–15, 2026) — Amsterdam, Netherlands. Annual gathering of the Bazel community uniting build system engineers, maintainers, and ecosystem contributors around fast, reproducible, multi-language software builds.
All Things Open 2026 (October 19–20, 2026) — Raleigh, North Carolina, USA. One of the largest community-focused open source conferences on the U.S. East Coast exploring open source software, AI engineering, and maintainer sustainability—featuring a Google keynote on Generative UI and the Open Future alongside the Google Community Lounge with the Flutter team.
PyTorch Conference North America 2026 (October 20–21, 2026) — San Jose, California, USA. Two days of open source machine learning innovation covering training, inference, compiler toolchains, and hardware heterogeneity across the PyTorch ecosystem.
AGNTCon + MCPCon North America 2026 (October 22–23, 2026) — San Jose, California, USA. Linux Foundation conference focused on building and scaling production agentic AI systems with open standards, observability, and control, featuring a keynote by Google's Rao Surapaneni.
GitHub Universe 2026 (October 28–29, 2026) — San Francisco, California, USA & Virtual. Annual global developer gathering highlighting open source workflows, collaborative security, and AI-assisted software engineering.
🗓️ November 2026
Open Source in Finance Forum (OSFF) New York (November 4–5, 2026) — New York, New York, USA. Dedicated industry conference examining open source compliance, supply-chain security, and collaborative innovation across regulated financial institutions.
KubeCon + CloudNativeCon North America 2026 (November 9–12, 2026) — Salt Lake City, Utah, USA. The Cloud Native Computing Foundation's flagship conference uniting adopters and maintainers around Kubernetes, platform engineering, and cloud-native infrastructure.
SFSCON (South Tyrol Free Software Conference) 2026 (November 13–14, 2026) — Bolzano, Italy. One of Europe's longest-running Free Software conferences bringing together public-sector decision-makers and developers to advance digital sovereignty and open infrastructure.
Open Source Reads and Links
[Article] Mila and Mozilla announce new initiative to build trustworthy open source AI for everyone, with Canadian government support — As a Canadienthusiast, seeing Mozilla and Mila partner with the Canadian government on an open source AI foundation layer immediately grabbed my attention. The crossover of open source and public-sector sovereign systems seems clear on the surface, yet real-world adoption reveals unexpected decisions. How do we make open source and open models useful to governments—is it in the open source infrastructure around them?
[Article] Google’s open source EnvHarness lets AI agents train against environments that evolve with them — As AI adoption matures, the need for solid infrastructure around AI is becoming more evident. It is nice to see Google Research’s work leading to open source tools like EnvHarness, which adapts training sandboxes to an agent’s failure modes—and raises the question of how our evaluation infrastructure must evolve alongside the agents themselves.
[Blog] Inside LLM Inference: Every Calculation from Text to Token using Gemma 4 12B — LLM inference often feels like an opaque black box, making Amulya Bhatia’s calculation-by-calculation trace of a token moving through Google’s open source Gemma 4 12B fascinating to ponder. Seeing how hybrid attention bounds memory under the hood prompts a broader question: how differently do we design AI systems when we actually understand the math happening inside the model?
[Video] 100% Local RAG Without Internet: Qdrant Edge and Google LiteRT — GDE Tarun R Jain pairs Google LiteRT and Gemma 4 with Qdrant Edge to run retrieval-augmented generation 100% offline. It prompts an interesting question for system design: how differently do teams experiment with their knowledge bases when local prototyping is free—and how much of what we default to the cloud could stay on the edge?
[Post] UISurf: An Operator-Centric Multi-Agent Platform for Observable and Cross-Environment UI Automation — UISurf explores orchestrating UI agents across Web, Desktop, and Mobile boundaries using the open Agent2Agent (A2A) protocol. Beyond the tool itself, its lessons on sandboxing and human-in-the-loop oversight are worth pondering: when autonomous agents coordinate across environments, what level of observability do human operators actually need to stay in control?
[Article] PS5 Linux lead quits as open-source projects have become “a bunch of noobs using LLMs” that “they don’t even understand” — When maintainers leave high-profile projects like PS5 Linux, it is important to look at the reasons. In the age of AI-assisted software development, unvetted LLM code and shifting bounty incentives are straining both maintainers and community trust. What can we learn from these events as we update open source and AI governance best practices to keep projects safe, secure, and viable?
Which of these stories will you be chatting about at your next meetup or conference? Let us know! Share with us on our @GoogleOSS X account or our @opensource.google Bluesky account.
The Apache Iceberg community recently merged a change to the REST Catalog specification that gives lakehouses something they have long been missing: a standard, engine-neutral way to express column masking and row filtering. It is a small change, but it addresses a structural gap in how open table formats handle data governance. This post walks through the problem, the design, and why the details matter.
The problem: there is no server in the read path
In a traditional database, access control is straightforward because there is exactly one door. Every query passes through the database server, the server knows who is asking, and if policy says "this user only sees the last four digits of the card number," the server applies that policy before returning results. One process, one enforcement point.
Apache Iceberg deliberately removed that door. The data is Parquet files in object storage, and any engine—such as BigQuery, Spark, Trino, Flink, PyIceberg, or DuckDB—can read those files directly. This design is what makes Iceberg fast and interoperable. But it also means that once an engine asks the catalog "where is the payments table?" and receives the metadata pointer, it has full access to the files. The catalog can grant or deny access to the whole table, but it has had no vocabulary for "access granted, but mask the email column" or "access granted, but only on rows where region = 'US'."
Until now, Iceberg had no built-in answer to this. The REST spec did offer coarser instruments like credential vending, which controls storage access per table, and server-side scan planning, which lets a catalog withhold entire data files. But neither can express "mask this column" or filter rows that are not already physically separated into their own files. So users who needed fine-grained access control had to step outside the open protocol entirely. They could adopt a vendor-provided client that understands its proprietary policy format, or route every read through a vendor's proxy service, giving up the direct-to-storage performance that motivated the use of Iceberg in the first place. It also quietly undermines Iceberg's core promise: the moment governance requires a specific vendor's client, the table is no longer open to any engine.
The new read-restrictions field in the REST Catalog spec is the community standardizing that vocabulary.
The mechanism
When an engine calls loadTable, the response may now include an optional read-restrictions object:
Read restrictions are expressed in two fields:
required-row-filter: a standard Iceberg predicate expression. Rows for which it evaluates to false must not appear in the result, and no information derived from them may be included.
required-column-projections: a list of columns, identified by field ID, each with a transformation the reader must apply before returning values.
One evaluation rule ties them together: the row filter is evaluated against the original, untransformed column values, and projections are applied to the rows that survive. This ordering is what makes the two features composable. A policy can filter on region = 'US' and also mask region in the output, and the filter still works. If masking ran first, any policy that filtered on a masked column would silently break.
Note what the catalog is doing here. It evaluates the access policy server-side—it knows the caller's identity from the authentication token—and returns only the result of that evaluation. The policy itself, with its roles, tags, and governance model, never crosses the wire. The engine does not need to understand how any particular catalog models governance. It needs to understand nine actions and a predicate.
The whole architecture fits in one sequence:
The catalog stays the policy decision point; the trusted engine becomes the policy enforcement point; the end user never holds storage credentials. Sections below unpack the three load-bearing details in this picture: the enforcement ordering, the fail-closed branch, and the trust boundary.
The nine masking actions
The specification defines a closed set of nine masking actions, each with an exact, per-type definition. The goal is cross-engine consistency: Spark, Trino, and PyIceberg must produce identical output for any given masking action. These actions are being implemented in iceberg-core, providing engines with spec-compliant transformations out of the box rather than requiring them to reimplement byte-level logic independently.
Below, the actions are grouped by the analytical utility of the resulting masked data:
Preserve the shape, hide the value
mask-alphanum—digits become n, other characters become x, with a small allowlist of punctuation (( ) , . - @) preserved. iceberg16112018@apache.org becomes xxxxxxxnnnnnnnn@xxxxxx.xxx—recognizably an email address, but not whose.
show-first-4 / show-last-4—preserve four code points, and apply mask-alphanum to the rest. 4111-1111-1111-4444 becomes nnnn-nnnn-nnnn-4444, the familiar customer-support view of a card number.
Hide everything
replace-with-null—the value becomes NULL. Only valid for optional fields; a server must not return it for a required field, and a reader that receives one must fail the query.
mask-to-fixed-value—the value becomes a type-specific constant (0, "XXXXXXXX", the epoch, an all-zero UUID, an empty list, and so on, each spelled out in the spec). Uniquely among the actions, this one also replaces NULL inputs—even the null-or-not bit is hidden.
sha-256-global—deterministic SHA-256, with exact byte-encoding rules per input type. The same input always produces the same output, everywhere, so GROUP BY user_id and joins across tables on a hashed key still work. The cost of that determinism is that hashed values can be tested against precomputed guesses—this is pseudonymization, not encryption.
sha-256-query-local—the same hash, salted with a fresh, cryptographically random salt (at least 16 bytes) per query. Values remain consistent within a single query, so self-joins and aggregation work, but cannot be correlated across queries, and precomputed-guess attacks no longer apply.
This last pair is a nice piece of design: the tradeoff between linkability and privacy, expressed as two enum values that a policy author chooses between per column.
Every action produces a value of the same type as its input—masked strings are strings, truncated dates are dates—so restrictions never change the schema an engine plans against. Queries do not need rewriting; values simply arrive transformed.
Fail-closed by design
The most consequential sentence in the specification is this one:
If a trusted reader that supports read-restrictions cannot apply any returned restriction, it must fail the query and must not silently return raw, partial, or empty results.
Consider the failure modes. A catalog sends an action added in a future spec version that the engine does not recognize. Or an expression type it cannot evaluate. The convenient behavior would be to skip what it does not understand and return the data. The spec rules this out: unrecognized action, fail; unparseable filter, fail; duplicate field ID in the projections, fail. Every ambiguity resolves to "no data" rather than "raw data." This is the right default for an access control mechanism, and it is also the one implementers would be tempted to soften—which is exactly why it is normative in the spec rather than left to judgment. It is also what allows the action vocabulary to grow in future versions without older engines becoming silent leak vectors.
A few prohibitions in the spec reward a closer look, because each one closes a subtle correctness hole:
No projections on map keys. Masking a map's keys can collapse two keys into one, or produce null keys, which engines silently coalesce or reject—data corruption presented as privacy. The spec bans it outright.
No projection on both a nested type and a field inside it. Masking a struct and also a field within that struct has no well-defined order of operations, so the spec refuses to define one: servers must not send it, and readers must reject it.
Everything references field IDs, never column names. This is standard Iceberg discipline: if ssn is renamed to national_id, the policy remains bound to the same physical column. A name-based policy would silently detach on rename—the worst possible failure mode for access control.
The trust model, stated plainly
All of this is enforced by the reader. The catalog hands the engine the file locations along with the restrictions, and a client that chooses to ignore the restrictions can read the raw files. So what is this actually protecting?
The specification is explicit: this mechanism assumes a trust relationship between the catalog and the engine, and how that trust is established is deliberately out of scope. The intended deployment is one where a platform team's engines—the shared Spark and Trino clusters—are trusted enforcement points that hold storage access (for example, through credential vending), while end users only ever talk to those engines and never hold storage credentials themselves. The trusted engine becomes the enforcement point, playing the role the database server played in the traditional architecture. The difference is that its behavior is now defined by a common, open protocol rather than by N proprietary integrations.
In other words, read restrictions do not protect data from the engine; they let the catalog direct a trusted engine to protect data from the engine's users. For genuinely untrusted readers, the coarse-grained model still applies: they are restricted to vended credentials where the storage access granted to the user aligns with the data permission of the user.
One operational subtlety deserves attention: restrictions are per-response and per-identity. The same loadTable call made by a different principal—or by the same principal later—may return different restrictions, and the restrictions attach to every read performed with that response, including subsequent planTableScan and fetchScanTasks calls. The spec therefore requires that the response not be cached outside its authentication scope. If your platform caches loadTableResponse, that cache is now security-sensitive and worth an audit.
What is still missing
Read restrictions are a foundation, not the finished building. It is worth being honest about the gaps between this specification and complete fine-grained access control:
The action vocabulary is fixed, because Iceberg Expressions are not implemented yet. Nine actions cover the common masking policies, but they are a deliberately closed set: a policy like "apply this custom redaction function" cannot be expressed, and the row filter is limited to predicates—comparisons that produce true or false. The path to generalizing this already exists on paper: the Iceberg Expressions specification, proposed by Ryan Blue and adopted in mid-2026, defines a portable structure for value expressions—constants, field references, and calls to well-defined functions or SQL UDFs. Once engines can evaluate those expressions, a catalog could return arbitrary transformations instead of choosing from an enum. Today no engine implements general expression evaluation, which is why the initial design confines itself to a small vocabulary that can be specified bit-for-bit.
There is no policy definition—deliberately. The specification standardizes the result of policy evaluation, never the policy itself. How an organization expresses "analysts see masked PII, auditors see everything"—the roles, tags, rules, and administrative APIs—remains entirely the catalog vendor's domain, and the assumption is that it stays there. Only the consequences of a policy are portable across engines; the policy definition is not. Whether communities eventually want a portable policy format is an open question the spec does not attempt to answer.
Trusted clients are asserted, not proven. As the trust-model section noted, how a catalog establishes that a caller is a trusted, enforcing engine is out of scope. In practice that trust is deployment configuration—service identities, network boundaries, and which principals receive vended credentials. There is no attestation mechanism in the protocol by which an engine proves it enforces restrictions; the trusted-client mechanism is an assumption the platform operator must make true.
None of these gaps undermines the design—each is a deliberate scoping decision that kept the proposal small enough to reach consensus—but they define the roadmap for what "complete" fine-grained access control in the open lakehouse still requires.
Why this matters
The specification change itself does not ship enforcement; that work in the engines begins now, starting with the default actions. But the shape of the design is right in three ways:
It picks the honest enforcement point. In an architecture with no server in the read path, the trusted engine is the only place enforcement can live without giving up direct storage reads. The spec accepts that constraint explicitly rather than obscuring it.
It standardizes the narrow waist. Catalogs keep their own rich policy engines—roles, tags, and attribute-based rules. Engines implement nine actions and a predicate evaluator, once. N×M becomes N+M.
It is fail-closed everywhere, which is what makes the vocabulary safely extensible.
The pattern—the catalog evaluates policy against the caller's identity and returns a small, closed vocabulary of obligations that the client must enforce or fail—is a useful template, and it would not be surprising to see more of the governance surface expressed this way over time.
The proposal was approved on the Apache Iceberg dev list with eight binding +1 votes and no objections. This is a strong signal of consensus across the many companies and open source communities that participate in the project. Consistent, engine-independent enforcement of fine-grained policies is a property the ecosystem has long wanted; it now has a specification for it, and the interesting work of implementing it in engines and catalogs is underway. If you work on either, the dev list is the place to get involved.
Did you know the Google Chrome Wikipedia page received over 36 million page views in 2025, or that Chromium logged over 3,700 commits in April 2010? If you have needed large, real-world datasets to benchmark query engines on Apache Iceberg, Google Cloud’s Lakehouse team is excited to announce the release of new public datasets in Google Cloud Lakehouse to help you explore and analyze open data at scale.
Google Cloud’s Lakehouse provides a high-performance storage catalog using Apache Iceberg as its open table format. By decoupling storage from compute, it enables you to use your preferred query engines—such as BigQuery, Apache Spark, or Trino—while managing data in an open format to avoid vendor lock-in.
These new public datasets are designed to help you explore the Apache Iceberg ecosystem and begin working immediately with real-world data. We are providing access to some of the most popular BigQuery public datasets, including Wikipedia pageviews and GitHub commit histories. Let’s dive into how you can start querying them.
Prerequisites
A Google Cloud project (for authentication).
Standard Google Application Default Credentials (ADC) configured in your environment.
Explore metadata with PyIceberg
To explore table metadata using PyIceberg, install the required Python packages in a virtual environment:
After installing the required packages, you can inspect a table’s schema using the Python script below. Here is how to inspect the table containing Wikipedia page view data from 2016:
Note: Replace <YOUR_PROJECT_ID> with your actual Google Cloud Project ID. This is required for the REST catalog to authenticate your quota usage, even for free public access.
PySpark and Managed Service for Apache Spark
Now that we have explored the table’s schema, we can use a managed, serverless Spark notebook to query the tables. This provides the speed and flexibility of Apache Spark without creating or managing a cluster. You can run this script in Google Cloud's serverless Managed Service for Apache Spark (formerly Dataproc):
With the session active, you can query monthly Wikipedia view counts for BigQuery-related articles by executing the following code in a new cell:
df1 = spark.sql("""
SELECT
title,
wiki,
SUM(views) AS total_views
FROM wikipedia.pageviews_2026
WHERE datehour >= TIMESTAMP '2026-01-01 00:00:00'
AND datehour < TIMESTAMP '2026-02-01 00:00:00'
AND LOWER(title) LIKE '%bigquery%'
GROUP BY title, wiki
ORDER BY total_views DESC
LIMIT 20
""")
df1.show(10)
Or, you can retrieve recent GitHub commits referencing Iceberg using the following snippet:
df2 = spark.sql("""
SELECT
commits.commit,
commits.subject,
commits.message,
commits.author.name AS author_name,
timestamp_seconds(commits.committer.date.seconds) AS commit_time,
repo_name
FROM github_repos.commits AS commits
WHERE LOWER(commits.subject) LIKE '%iceberg%'
OR LOWER(commits.message) LIKE '%iceberg%'
ORDER BY commit_time DESC
LIMIT 50
""")
df2.show(10)
Start building today
These datasets were imported from their BigQuery counterparts and transformed using Apache Iceberg as the table format and Parquet as the data file format. Our goal is to lower the entry barrier so you can learn and explore Apache Iceberg with your favorite query engine without managing any infrastructure. To get started with building an open, managed, and high-performance Iceberg lakehouse, visit the Google Cloud Lakehouse page.
Open source underpins an $8.8 trillion global economy, yet sustaining it requires moving beyond voluntary charity and philosophical appeals. This week’s reads examine how the ecosystem is adapting to modern economic and AI pressures—from proposing package-registry royalties and quantifying 2–5x net value for regulated utility grids, to turning autonomous AI security incidents into hardened supply-chain defenses that empower open-science breakthroughs like NASA and IBM’s Lunar Foundation Model.
Upcoming Events
🗓️ October 2026
MCP Dev Summit Toronto (October 5–6, 2026) — Toronto, Ontario, Canada. Dedicated Linux Foundation developer summit advancing the Model Context Protocol (MCP) and open agentic interoperability standards. Join Google OSPO's Daryl Ducharme on October 6 for "Tag-Team Transmission: Navigating A2A and MCP for Optimum Orchestration," examining how the Agent2Agent (A2A) protocol and MCP interoperate across multi-agent architectures.
Open Source Summit Europe (October 7–9, 2026) — Prague, Czechia. The premier Linux Foundation gathering in Europe celebrating the 35th anniversary of Linux and connecting developers, technologists, and community leaders across open AI, embedded systems, and digital trust.
Community Over Code (October 11–14, 2026) — Glasgow, Scotland. The flagship Apache Software Foundation conference bringing together project maintainers, committers, and users to collaborate on open governance, data architecture, and community-led software development.
All Things Open 2026 (October 19–20, 2026) — Raleigh, North Carolina, USA. One of the largest community-focused open source conferences on the U.S. East Coast exploring open source software, AI engineering, DevSecOps, and maintainer sustainability.
GitHub Universe 2026 (October 28–29, 2026) — San Francisco, California, USA & Virtual. Annual global developer gathering highlighting open source workflows, collaborative security, and AI-assisted software engineering.
🗓️ November 2026
Open Source in Finance Forum (OSFF) New York (November 4–5, 2026) — New York, New York, USA. Dedicated industry conference examining open source compliance, supply-chain security, and collaborative innovation across regulated financial institutions.
KubeCon + CloudNativeCon North America 2026 (November 9–12, 2026) — Salt Lake City, Utah, USA. The Cloud Native Computing Foundation's flagship conference uniting adopters and maintainers around Kubernetes, platform engineering, and modern cloud infrastructure.
SFSCON (South Tyrol Free Software Conference) 2026 (November 13–14, 2026) — Bolzano, Italy. One of Europe's longest-running Free Software conferences bringing together public-sector decision-makers and developers to advance digital sovereignty and open infrastructure.
Open Source Reads and Links
[Blog] Nobody pays for open source. We can force them to. — Companies already spend over $1 billion a year on open source, but that money goes to supply-chain mirror and security vendors rather than the maintainers writing the code. Laurie Voss breaks down why voluntary charity and license changes always fail, arguing that commercial package registries and mirror vendors should pay automatic royalties down the dependency tree to long-tail maintainers.
[Article] AI Is reshaping open source software and straining the systems that sustain it — As conversations grow around moderating the speed of AI development across the industry, it is critical to watch the friction points where AI and open source overlap—particularly as an influx of AI-generated code and vulnerability discovery strains maintainers of an $8.8 trillion ecosystem. Funding remains front and center, requiring not just raw capital but continuous monitoring for effectiveness in sustaining the human governance, consensus-building, and non-coding infrastructure that AI cannot replace.
[Report] LF Energy Research Finds Open Source Software Can Deliver 2-5x Greater Net Value for Grid Operators — While open source software and public utilities seem like natural philosophical allies as public goods, philosophical appeals routinely fail in regulated infrastructure sectors where operators are bound by strict reliability mandates and ratepayer accountability. This new LF Energy report demonstrates how to bridge that gap: by translating collaborative "make together" governance into an auditable benefit-cost framework showing two to five times greater net value over proprietary procurement.
[Article] AI Is at a Turning Point — High-profile incidents of autonomous AI agents breaching Hugging Face or social-engineering open source maintainers inevitably dominate headlines and amplify purely negative narratives around AI. However, treating these security failures as concrete post-mortems—exposing how standard reinforcement learning incentivizes deceptive subgoals—is exactly what allows open source communities to establish hardened supply-chain defenses and safely advance high-impact, positive AI use cases.
[Article] IBM and NASA Release Open-Source AI Model to Support Lunar Exploration — Seeing open source, AI, open data, and space science converge is genuinely inspiring, especially when it demonstrates that open scientific models are only as useful as the open datasets beneath them. While NASA and JAXA planetary archives have long been public, IBM and NASA's release of the Lunar Foundation Model alongside the first harmonized, machine-learning-ready lunar dataset (unifying 30+ multi-instrument layers) proves that transforming raw open data into shared, interoperable infrastructure is what unlocks domain discovery.
Which of these stories will you be chatting about at your next meetup or conference? Let us know! Share with us on our @GoogleOSS X account or our @opensource.google Bluesky account.