OneLake learns Iceberg: what the Snowflake deal means for your data architecture
At Build 2024 Microsoft announced future Apache Iceberg support in OneLake and deeper interoperability with Snowflake. What was announced, what is still coming, and how to plan for it.
The table format war just got a little less important. At Microsoft Build today, Microsoft announced an expanded partnership with Snowflake and future support for Apache Iceberg in Fabric OneLake. The promise is that Fabric and Snowflake can work on a single copy of your data, instead of each platform keeping its own.
If your organisation runs both platforms, or is choosing between them, this changes the conversation. The question moves from "which platform owns the data" to "which engine is best for this job".
Most of it is not available yet. Microsoft describes it as future and upcoming. So here is what was said, what is actually in preview today, and how I would plan around it.
What was announced
Arun Ulag's Build blog post groups this under the Snowflake partnership. Microsoft says that since the launch of Fabric it has committed to open data formats, standards and interoperability with partners, and this is the next step.
The key points:
- Future support for Apache Iceberg in OneLake, next to Delta Lake, which Fabric already uses.
- Upcoming shortcuts for Iceberg in OneLake, so Fabric users can reach Iceberg data, including data written by Snowflake.
- Translation of metadata between the Iceberg and Delta formats.
- Analysis of Fabric and Snowflake data written in Iceberg format in any engine within either platform.
- OneLake data can be accessed in Snowflake as well as in Fabric.
The headline line is that you can work with a single copy of your data across Snowflake and Fabric. No dates were given, and no preview status was stated for the Iceberg parts.
Why the table format matters
A table format is the layer on top of plain Parquet files that turns a folder into a table. It holds the schema, the list of files and the transaction history. Delta Lake and Apache Iceberg solve the same problem in different ways, and for years they have been the dividing line between ecosystems.
Fabric is built on Delta, and this announcement meets Snowflake on Iceberg. Until now, getting data from one side to the other has meant a copy: an export, a pipeline, a second storage bill and a second version of the truth that can drift.
Metadata translation is the interesting idea here. The Parquet files can stay where they are, and the platforms only need to agree on how to describe them. If that works as described, the cost of having two engines drops from "duplicate the data" to "maintain a shortcut".
Shortcuts are the delivery mechanism
Shortcuts are already the way OneLake connects to data it does not own. A shortcut is an object in OneLake that points to another storage location, a bit like a symbolic link. It appears as a folder, and Fabric engines can read through it. When you create a shortcut in the Tables section of a lakehouse and the target is in Delta format, the lakehouse recognises it as a table.
Iceberg shortcuts extend that pattern to a second format. The Fabric team says it plans to add support for Iceberg tables in OneLake shortcuts, so table shortcuts will no longer be limited to Delta Lake data.
Something you can try today: Microsoft also announced a preview of OneLake shortcuts to on-premises and network-restricted data sources. With the on-premises data gateway, you can create shortcuts to Amazon S3, S3 compatible storage and Google Cloud Storage that sit behind a firewall. The guidance is the same as for any shortcut: if the data is already in Delta format, put the shortcut in Tables, otherwise in Files.
Not only Snowflake
The same post announced a deeper integration with Azure Databricks, also marked as coming soon. You will be able to access Azure Databricks Unity Catalog tables directly in Fabric, and Fabric items such as lakehouses will be available as a catalog in Azure Databricks, with the data staying in OneLake.
Put together, the direction is clear. Microsoft is positioning OneLake as the storage layer that other engines can read and write, rather than a closed store that only Fabric uses.
What this means for you
Do not change your architecture on an announcement. Do prepare for it. After two decades in data, I have seen many "single copy" promises meet the reality of permissions, performance and version mismatches.
Here is what I would do now:
- Map where your data is copied between platforms today. Every pipeline that only moves data from Snowflake to Fabric, or the other way, is a candidate to remove later.
- Standardise on open table formats for new work. Write Delta in Fabric, and if you are on Snowflake, look at Iceberg tables for data that other engines need to read.
- Get comfortable with shortcuts now, using Delta data you already have. The governance questions (who owns the target, who can read through the shortcut) are the same whatever the format.
- If you have data behind a firewall in S3 compatible storage, test the on-premises shortcut preview on a non-critical dataset.
- Keep decisions reversible. Avoid building new copy processes that will be hard to retire if the Iceberg support lands as promised.
Two things to watch when it does arrive. First, which direction works on day one, since read access and write access are rarely released together. Second, how security and performance behave when one engine reads tables managed by another.
Takeaway
The real news is not a single feature. It is that Microsoft and Snowflake are both saying the storage layer should be open, and that the engine should be a choice you make per workload. That is good for anyone designing a data platform, because it lowers the cost of being wrong.
Are you running Fabric and Snowflake side by side today, and how many copies of the same data are you paying for?
Sources
Enjoyed this? Get the next one by email
Occasional emails about Microsoft Fabric, SQL Server, Power BI and Synapse.