Google Cloud has introduced a borderless Lakehouse for querying and governing data spread across Google Cloud, Amazon Web Services, Microsoft Azure, Databricks, Snowflake and on-premises systems. It combines BigQuery, the open Apache Iceberg format and Knowledge Catalog to give AI agents common context without immediately copying every dataset into one warehouse.
The proposal addresses a real problem. An agent may have an excellent model and still make poor decisions because sales, inventory, contracts and technical logs live in separate silos. Yet borderless does not mean transfer-free, cost-free or independent from Google. Each query path still needs architectural scrutiny.
The short answer
| Question | Answer |
|---|---|
| What is Google trying to unify? | Discovery, governance and analysis across several clouds and data platforms. |
| Must all data migrate? | No. Some data can be queried remotely; other data will be replicated or materialized for performance. |
| Why Apache Iceberg? | It separates table metadata from data files and improves interoperability between engines. |
| What does Knowledge Catalog do? | It indexes assets, business meaning, quality, lineage and access rules. |
| Does that automatically make data agent-ready? | No. Reliable definitions and permissions are still required to prevent false or excessive context. |
| What is the main risk? | Treating the layer as universal while latency, network charges, permissions and features vary by source. |
The challenge is context, not merely storage
An enterprise may keep transactions in BigQuery, historical files in Amazon S3, an operational database on Azure and analytics notebooks in Databricks. Dashboards can tolerate some separation. An autonomous agent must discover the correct data, understand its freshness, obtain permission and act before its context becomes stale.
Copying every source into a central repository simplifies some analytics while creating delays, duplication and new ownership duties. Querying each system remotely avoids migration but produces less predictable latency and availability. The borderless Lakehouse tries to support both approaches instead of imposing one.
This distinction matters. A one-off query against a small dataset can remain remote. A calculation repeated by thousands of agents will probably need caching, materialization or compute placed nearer the data. The logical view may look unified while bytes remain subject to physical networks and egress bills.
Apache Iceberg provides an open table contract
Iceberg describes the schemas, partitions, snapshots and files composing an analytical table. Multiple engines can therefore work with Parquet data without adopting a separate proprietary table format. Google positions its managed Iceberg tables as an open lakehouse foundation stored in customer-controlled buckets.
Openness still has rules. BigQuery documentation warns that directly modifying files tracked by a managed Iceberg table can make it inconsistent. Overlapping storage prefixes can even cause automated garbage collection to delete files treated as untracked. An open format does not mean every engine can write simultaneously without coordination.
Organizations need a clear write owner, catalog management and tested maintenance procedures. Interoperability is meaningful for reading and selected workflows, but it does not eliminate the operational controls of a managed service.
BigQuery Omni moves computation closer to data
BigQuery Omni runs BigQuery analytics against Amazon S3 or Azure Blob Storage through BigLake tables. Google documents several strategies: remote joins, materialized views, query-based transfers and full loads. Each balances freshness, cost and performance differently.
A remote join can suit occasional exploration at a reasonable scale. A materialized view is more appropriate for a dashboard or agent repeatedly asking the same family of questions. Transfers become attractive when transformations are complex or network round trips are too expensive.
Claims about querying data in place must therefore be measured. Teams should observe bytes scanned, cross-region movement, startup time and SQL operator limitations. A prototype using a few gigabytes cannot predict the cost of an agent fleet continuously reading terabytes.
Knowledge Catalog becomes the map agents follow
Google has renamed Dataplex Universal Catalog to Knowledge Catalog and describes it as a Gemini-powered catalog. It indexes technical and business metadata, searches assets through natural language and builds context for analysts and agents.
This layer matters as much as query execution. Two columns named revenue may represent gross sales in one system and recognized revenue in another. Without definitions, owners, refresh frequency and quality indicators, an agent can join technically compatible but semantically incorrect data.
Automated descriptions help bootstrap a catalog; they cannot invent governance. Domain teams must validate important terms, sensitive classifications and retention rules. A wrong context graph can propagate one error through every response while making it look consistent.
Permissions must survive unification
A common interface must not flatten access rights. An inventory agent does not need payroll records even if both appear in one catalog. Service identities, column policies, masking and audit logs should be enforced at every access.
The problem grows when Google policies coexist with AWS IAM, Azure roles and third-party permissions. A lowest-common-denominator mapping may be too broad, while strict translation can block legitimate work. Automated authorization tests need both allowed and denied scenarios.
Agents add another hazard because they generate queries and select tools dynamically. A hostile instruction hidden in a document must not expand access. Permissions, source allowlists, budgets and approvals for sensitive actions should be imposed outside the model.
Multicloud remains distributed computing
Results depend on the catalog, query engine, connector, remote cloud and network. One failing layer can leave an agent partially blind. Responses should carry source freshness and stop safely when mandatory data is unavailable.
Observability must include the generated query, tables accessed, schema version, bytes read, policy applied, intermediate result and final decision. Without that trace, an organization cannot explain why an agent recommended replenishment or blocked a transaction.
Recovery objectives should also distinguish catalog metadata from data files. Restoring files without their metadata may not recreate a usable table. Restoring a catalog that points to deleted objects provides an accurate map of a territory that no longer exists.
How to evaluate the borderless Lakehouse
A pilot should begin with one workflow, such as reconciling AWS inventory with BigQuery orders to recommend replenishment. Teams should measure accuracy, freshness, cost per decision, latency and behavior when one source is unavailable.
They can then compare remote queries, targeted replication and full centralization. The practical architecture will often be hybrid. Large stable datasets remain close to storage, frequently used aggregates move nearer compute, and sensitive information retains local controls.
The borderless Lakehouse is therefore not the disappearance of boundaries. It is a translation and governance layer above them. Its value will depend on catalog quality, disciplined Iceberg writes and transparent multicloud costs. For AI agents, accessible data only becomes useful when it is also understood, authorized, fresh and traceable.




Join the discussion
Comments
Loading comments…