Fabric Can Finally Monitor Itself: Building a Real-Time Capacity Control Loop

Most Fabric capacity incidents are diagnosed in reverse.

A report slows down. A refresh misses its window. A user sees a capacity-limit error. Then someone opens the Capacity Metrics app and works backward through the evidence to find out what happened.

That is useful monitoring, but it is still an autopsy.

Capacity Overview Events change the sequence. Microsoft made them generally available in this month, giving Fabric a near-real-time stream of each capacity’s smoothed utilization and throttling posture. A Summary event represents a 30-second window for an active capacity, except that all-zero windows are suppressed. State events arrive when status changes. Those signals can flow through Real-Time hub and Eventstream, land in Eventhouse, appear on a Real-Time Dashboard, trigger Fabric Activator, and ultimately drive a guarded scaling action through Azure Resource Manager.

This is also a useful way to reintroduce Real-Time Intelligence. RTI is sometimes described as Fabric’s specialist workload for telemetry, clickstreams, and IoT. Capacity Overview Events show the larger idea. Fabric can use its own event-driven architecture to observe and operate Fabric.

The goal is not “add an alert when utilization reaches 80 percent.” It is an architecture that turns capacity management into a controlled feedback loop.

Continue reading “Fabric Can Finally Monitor Itself: Building a Real-Time Capacity Control Loop”

Before the Capacity Fire Starts: Why FUAM Belongs in Every FSS Fabric Baseline

Most Fabric monitoring conversations begin too late.

They begin when a workspace is already noisy, when refreshes are already failing, or when capacity pressure has already become visible enough to trigger concern. By that point, the real problem is usually larger than utilization. It is a visibility problem. Teams do not have a durable, tenant-level view of what exists, what changed, what is connected to source control, what is actively used, and where governance has started to drift. That is the gap FUAM was designed to close. The distinction worth making is not between a “simple” and a “serious” version of monitoring, but between an original governance-first package and the newer Fabric Toolbox implementation that extends that foundation with Capacity Metrics and deeper operational analysis.

That distinction matters for FSS environments because it changes the order of operations. The older GT-Analytics package, labeled FUAM Basic, focused on the data that can be gathered without reading Capacity Metrics. The current Fabric Toolbox implementation keeps the same broad monitoring vision, but adds Capacity Metrics as part of a larger platform-admin monitoring solution. If the question is what should be present by default in an FSS Fabric setup, the answer starts with the governance-and-inventory layer and then grows into the richer Capacity Metrics-enabled implementation as operational maturity increases.

Continue reading “Before the Capacity Fire Starts: Why FUAM Belongs in Every FSS Fabric Baseline”

FinOps for the Data + AI Era: Strong Structures Beat Strong Opinions

The fastest way to turn cloud enthusiasm into executive skepticism is simple: ship something impressive in Data or AI…and then hand Finance a bill no one can explain.

That’s not a tooling problem. It’s a structure problem.

In this post, I’m going to make the case for strong FinOps structures that actively engage Data, AI/ML, and the broader cloud stack—not as a “cost police” function, but as an operating model for technology value. We’ll look at why the scope of FinOps has expanded, what makes Data and AI spend uniquely tricky, and what “strong” actually looks like when it’s working.

Continue reading “FinOps for the Data + AI Era: Strong Structures Beat Strong Opinions”

From Telemetry to Trust: Using FUAM + Purview Lineage to Make Fabric Governance Pay Off

If you’re running Microsoft Fabric at any real scale, you’ve probably felt the tension: the platform makes it easy to build, share, and iterate—but it also makes it easy to spend, sprawl, and accidentally ship the wrong answer.

The good news is you already have most of the raw ingredients to fix that. What’s missing is an operating model that converts “platform signals” into business outcomes: predictable costs, cleaner estates, and faster response when data is wrong.

In this post I’ll walk through three practical patterns:

  • using FUAM as a telemetry backbone for FinOps that people will actually use
  • using the same signals for stale workspace detection (without manual audits)
  • combining Microsoft Purview lineage with usage signals to identify incorrect datasets that are actively being consumed—and contain the blast radius

Along the way, I’ll stay grounded in business value: what these ideas buy you in dollars, time, and trust.

Continue reading “From Telemetry to Trust: Using FUAM + Purview Lineage to Make Fabric Governance Pay Off”

Stop Paying Hot-Tier Prices for Cold Data: Using ADLS Gen2 to Tame Fabric Ingestion Storage Costs

If you’ve been living in Microsoft Fabric for a few months, you’ve probably felt it: the platform makes it incredibly easy to ingest data… and surprisingly easy to rack up storage spend while you’re doing it (especially considering how much storage is included).

The pattern is common. A team starts with a Lakehouse, adds Pipelines or Dataflows Gen2 for ingestion, follows a sensible medallion approach, and before long they’re keeping “just in case” raw files, repeated snapshots, and long-running history inside OneLake—often at the same performance tier as yesterday’s data. The storage bill grows quietly. Capacity pressure shows up in places you didn’t expect. And suddenly “simple ingestion” is a FinOps conversation.

Here’s the good news: you don’t have to choose between Fabric and sensible archival strategy. Azure Data Lake Storage Gen2 (ADLS Gen2) can be your pressure relief valve—your durable landing zone and archive—while Fabric stays the place you compute, curate, model, and serve.

What follows is a deep dive into how to use ADLS Gen2 accounts to solve the archival and storage-cost traps that show up during Fabric ingestion: where the costs come from, what architectural patterns work well, and the practical implementation details (shortcuts, security, and billing mechanics) that make it real for Microsoft Fabric teams.

Continue reading “Stop Paying Hot-Tier Prices for Cold Data: Using ADLS Gen2 to Tame Fabric Ingestion Storage Costs”

Real‑Time Data Isn’t Free: The Complexity and Cost Tradeoffs (From Trickle to Internet‑Class)

The first time someone asks for “real‑time,” it sounds like a small tweak: refresh the dashboard faster, trigger an alert sooner, show a counter that feels alive. In a data platform, that single request quietly changes everything—how you ingest, how you process, how you serve, and how you operate.

This post keeps it practical. It frames real‑time as a freshness target (not a vibe), walks through the two taxes real‑time introduces—architectural complexity and cost—and shows how patterns evolve as you scale from modest #StreamingData to internet‑class velocity. It also folds in recent Microsoft Ignite announcements that matter for real‑time platforms, including SQL Server 2025’s “change event streaming” and near real‑time analytics via OneLake/Fabric mirroring, plus the continued maturation of Microsoft Fabric’s Real‑Time Intelligence building blocks.

Continue reading “Real‑Time Data Isn’t Free: The Complexity and Cost Tradeoffs (From Trickle to Internet‑Class)”