Microsoft Fabric is here: a first take for Power BI and Synapse teams
Microsoft announced Fabric in preview at Build 2023. What it is, how OneLake and Direct Lake fit in, and what Power BI and Azure Synapse customers should do now.
Microsoft has just unveiled Microsoft Fabric at Build, and it is now in preview. Fabric brings Azure Data Factory, Azure Synapse Analytics and Power BI together into one software as a service (SaaS) product, built on a single data lake called OneLake.
If you run Power BI Premium, or a Synapse workspace in production, this announcement is about you. Here is what was announced, and what I would do with it.
What Fabric is
Fabric is a single analytics product with role-specific experiences on top of shared storage and shared compute. The announcement lists seven core workloads:
- Data Factory (preview), with more than 150 connectors and data pipelines.
- Synapse Data Engineering (preview), with Spark authoring and live pools.
- Synapse Data Science (preview), for building, training and managing models.
- Synapse Data Warehousing (preview), a converged lakehouse and warehouse experience.
- Synapse Real-Time Analytics (preview), for streaming and semi-structured data.
- Power BI, the reporting and BI experience most of us already know.
- Data Activator, which is coming soon and is in private preview.
The key word is SaaS. You do not provision storage accounts or Spark clusters. Copilot in Fabric is also promised, but marked as coming soon, so I leave it aside here.
OneLake is the real foundation
Microsoft calls OneLake "the OneDrive for data". Every Fabric tenant gets one automatically, and all the workloads are wired into it.
Three details matter more than the slogan:
- At the API layer, OneLake is built on and compatible with Azure Data Lake Storage Gen2. Existing tools that speak ADLS Gen2 should have a familiar surface.
- Delta on top of Parquet is the native, default format for all workloads. Load the data once, and Spark, the warehouse and Power BI can work on the same copy.
- Shortcuts let OneLake point at data in ADLS Gen2 and Amazon S3, with Google Storage listed as coming soon, without copying it.
For years, every engine has wanted its own copy of the data. A shared Delta layer is a different starting point, if it holds up in practice.
One thing to note: the universal security model managed in OneLake, enforced across all engines at table, column and row level, is described as coming soon. Until it lands, plan your security per engine as you do today.
Direct Lake: the Power BI headline
For Power BI people, Direct Lake is the feature to understand. A Power BI dataset in Direct Lake mode reads Delta tables in OneLake directly, without importing them.
Microsoft says Direct Lake datasets get query performance on par with import mode, with the real-time nature of DirectQuery, and no refreshes to manage. That is a strong promise, aimed at exactly the trade-off we have fought with for years.
The status is worth reading carefully. Direct Lake is in preview for datasets on Lakehouses. For Warehouses it is in private preview, although it works when you use the SQL endpoint of a Lakehouse. So test it on Lakehouse data first, and measure it against your own import models before you believe the comparison.
What it means for Power BI customers
Microsoft is clear that Power BI keeps all the functionality it has today. What changes is the admin side:
- Fabric is controlled by a tenant setting in the Power BI admin portal. It is off by default, and you can enable it for specific users or security groups.
- If the admin takes no action, Fabric will be turned on by default for all Power BI tenants starting July 1.
- With Fabric enabled, Premium capacities can run the new workloads. Microsoft says the new experiences will not draw down usage from Premium capacity before October 1, 2023, and you can watch the impact in the Capacity Metrics app.
- Power BI Pro users can get access through a 60-day Fabric trial, if trials are enabled by the tenant admin.
The July 1 default is the date to put in the calendar.
What it means for Azure Synapse customers
This is the question I expect to hear most. Microsoft's answer is that Azure Synapse Analytics, Azure Data Factory and Azure Data Explorer will continue as enterprise-grade platform as a service (PaaS) offerings. Fabric is described as an evolution of those offerings that can connect to them, and customers can move into Fabric at their own pace.
Read that as: no forced migration today, but a clear signal of where new investment is going. If you run a dedicated SQL pool in production, keep running it, and keep an eye on the direction.
What to do next
A sober plan for the coming weeks:
- Decide your tenant setting before July 1. Enable Fabric for a small security group rather than everyone.
- Start a trial or use a test Premium capacity. Do not try it first on your production capacity.
- Build one Lakehouse with real but non-sensitive data, and put a Direct Lake dataset on top.
- Compare it with an existing import model on query times and workflow.
- For Synapse estates, list your pipelines, Spark jobs and SQL pools, and note which ones you would actually want in a SaaS model. That list is your future roadmap, not a migration project for this quarter.
Almost everything here is preview, so expect behaviour to change before general availability.
Takeaway
After two decades in data, I have seen several "one platform for everything" announcements. Fabric is different in one way: it starts from a shared open storage format, not a shared logo. Whether that makes life easier depends on how OneLake, Direct Lake and capacity management behave under real workloads.
My advice is to be curious, not hasty. Try it, measure it, and make the tenant decision on purpose. Which part of your current Power BI or Synapse setup would you test in Fabric first?
Sources
Enjoyed this? Get the next one by email
Occasional emails about Microsoft Fabric, SQL Server, Power BI and Synapse.