(+351) 21 24 10006  ·  info@bconcepts.pt
Carnaxide, Lisbon
Data observability in Microsoft Fabric: a practical guide
Data Engineering

Data observability in Microsoft Fabric: a practical guide

João Barros 01/10/2026 9 min

Without clear observability mechanisms, data do not fail silently: they erode decisions and trust at a pace only noticed when it’s too late.

Why data observability is decisive in Microsoft Fabric

Data projects today run on fast cycles: continuous ingestions, transformations in lakehouses and semantic models consumed by reports and applications. In Microsoft Fabric, that chain combines familiar components (Parquet storage, Spark, Power BI) with orchestrated services and an integrated surface (OneLake, Lakehouse, Notebooks, Data Factory/Jobs). That brings productivity gains, but also distinct friction points — partitions that stop being updated, schemas drifting after an API change, expired credentials blocking loads and pipelines failing silently after upstream changes.

Data observability in Microsoft Fabric: a practical guide

Observability is not just a set of technical checks. It’s the ability to answer, with instrumented data, three basic questions: what changed, when did it change and what is the impact on business lines. Without that ability, reports degrade and response tends to be reactive and slow. Correctly implemented, observability reduces downtime, decreases analytical rework and restores user trust in Power BI and other data consumers.

In a practical context: imagine an e-commerce operations dashboard used by 15 managers that makes replenishment decisions based on stock data. If the daily ingestion fails without an alarm, decisions keep being made until discrepancies become obvious — typically 24–72 hours later — causing avoidable losses in sales and logistics. Observability aims to detect that failure in 15–30 minutes, not days.

Essential metrics: what to measure and why

Not all metrics are equal. Prioritize four dimensions that are easy to translate into actionable rules: freshness, completeness, integrity/values and schema drift. These four cover the most common failure modes that affect report consumption and automated processes.

Concrete examples of metrics and practical thresholds:

  • Freshness: time since last ingestion per table/partition. SLA example: delta_minutes ≤ 30 for operational tables, ≤ 1440 for historical tables. This translates into checks that compare the partition’s timestamp_max with now().
  • Completeness: percentage of expected records per partition. For deterministic sources (e.g., 24 hourly feeds), use row_count ≥ 95% of expected; for variable sources, base on moving medians (e.g., row_count ≥ 70% of the median of the last 7 runs).
  • Integrity: percentage of null values in critical fields. A practical threshold: nulls_in_customer_id ≤ 0.1% for transactional tables; for sensitive foreign keys, any increase above 0.5% should trigger an alert.
  • Schema drift: schema fingerprint per table and alerts on changes to columns, types or presence/absence of fields. Immediate alert for removal of a column used in calculations or models.

These metrics serve to create actionable triggers. Numerical example: in a table with 5M rows, an increase in nulls in customer_id from 0.01% to 1.5% translates to 75,000 affected records — a clear sign of a CRM integration break that justifies immediate investigation. Additionally, tracking the rate of change of data per column (for example, entropy or variance of the distribution) helps detect subtle changes that counters do not catch.

Practical observability architecture in Fabric

An effective architecture combines instrumentation at ingestion, tracing in the Lakehouse and integrated visualization/alerts. This should be simple and scalable: record metrics in the Lakehouse itself, expose them via Power BI and integrate alerts with Azure Monitor and Teams. Key elements to consider:

  • Pipelines (ETL/ELT) with verification steps: before and after each critical transformation, execute queries that record row_count, min/max timestamps, null percentages and samples of values. Those results should be written to a metadata table.
  • Metadata index in the Lakehouse: health tables that store row_counts, checksums, timestamps and schema fingerprints per partition. Build these tables partitioned by date and keep at least 90 days of history for trend analysis.
  • Storage of logs and metrics in an easy-to-query destination: Log Analytics for infrastructure telemetry and Parquet/Delta tables in OneLake for business metrics. Maintaining both sources allows combining execution logs (jobs) with data metrics.
  • Power BI dashboards that aggregate metrics by domain and alerts via Azure Monitor/Teams/Email for escalation. Include SLA KPIs, trending charts and a table of recent incidents with links to runbooks.

In Microsoft Fabric, use Jobs/pipelines to run checks at the end of each load and Spark Notebooks for heavier checks. Record results in metadata tables in the Lakehouse, which feed observability reports and triggers in Azure Monitor for real‑time notifications. A typical implementation consumes little additional space: 90 days of metrics for 200 tables may occupy 50–200 MB in Parquet, a marginal cost within OneLake.

Check implementations: light to heavy

There are low‑cost checks that should run on every ingestion and heavier checks that run on a less frequent schedule. The idea is to filter trivial problems quickly and reserve resources for deep analyses that require more I/O or CPU.

Light checks (run on every job):

  • Counts per partition and comparison with the expected value. Practical rule: if abs(cur - expected) / expected > 0.1 then alert. For expectations, use moving medians to avoid alerts due to seasonality.
  • Percentage of nulls in key columns and format validation (e.g., regex for NIF, timestamp is not null).
  • Freshness check with timestamp_max and send an alert if it exceeds the SLA.

Heavy checks (nightly or weekly):

  • Checksum of combined columns to detect value changes without count changes. E.g.: calculate md5(concat(col1, col2)) per partition and compare with history.
  • Reconcile volumes with source systems: compare daily billing totals between ERP and lakehouse; initial tolerance of 0.5% and investigate if it exceeds 1%.
  • Statistical drift analyses: compute KL divergence or KS test between current and historical distributions for numeric/categorical columns.

Practical approach: implement light checks as SQL inside the pipeline and write results to a _pipeline_health table. For heavy checks, run Spark notebooks in Fabric and keep a history of metrics for trend analysis. A simple example SQL query for a light check:

insert into lakehouse._pipeline_health
select
  dataset_name,
  partition_date,
  count(*) as row_count,
  max(event_ts) as last_ts,
  sum(case when customer_id is null then 1 else 0 end) * 1.0 / count(*) as null_ratio
from lakehouse.datasets.orders
where partition_date = current_date
group by dataset_name, partition_date

Storing these results allows building alerts that not only fire but also provide context for initial diagnosis.

Mini practical case: reducing incidents at a retail scale‑up

At a retail company with 80 employees and 5 main pipelines, daily loads brought 50 million rows into the lakehouse. Before observability, there were on average 8 incidents per month related to stock and pricing discrepancies, each taking on average 6 hours to detect and 10 hours to fix (involving devs, analysts and operations). Hidden costs came from labor hours and commercial impact.

Approximate monthly costs before the solution:

  • Lost hours: (8 incidents) × (16 hours per incident) × €60/hour ≈ €7,680
  • Estimated sales loss from pricing errors and stockouts: €12,000
  • Reputation and support costs: conservative estimate of €2,000
  • Total ≈ €21,680/month

Interventions implemented in 6 weeks:

  • Light checks in each pipeline (row counts, nulls, freshness) with thresholds based on medians of the last 30 runs.
  • Weekly schema fingerprinting and nightly checksum per partition to detect subtle value changes.
  • Health dashboards in Power BI and critical alerts via Teams integrated with Azure Monitor; documented runbooks for each alert type.

Results in 3 months:

  • Monthly incidents reduced from 8 to 1 — an 87% reduction.
  • Average time to detect issues fell from 6 hours to 30 minutes; average fix time fell from 10 hours to 2 hours.
  • Lost hours: (1 incident) × (1.5 hours detection + 2 hours fix) × €60 ≈ €210/month.
  • Sales loss from errors practically eliminated; monthly avoided loss estimate: €11,000.

ROI: initial implementation investment ≈ €18,000 (engineering, pipeline configuration and dashboard creation), payback in 1–2 months by reducing operational costs and lost sales. These figures illustrate that observability is a pragmatic investment: reducing detection and resolution time has direct impact on margins and user trust.

Observability means that, when a report stops being reliable, the team knows exactly where to start looking. This shortens the path from suspicion to resolution.

Practical alerts and how to avoid alert fatigue

Sending alerts for everything is a recipe for ignored alarms. To maximize effectiveness, categorize alerts by severity and apply noise reduction tactics:

  • Critical alerts: freshness failure in operational tables, loss of integrations with vendors or schema changes that break models — immediate notifications via Teams/Email with a clear runbook and designated owner.
  • Warning alerts: 5–10% deviations in counts, moderate increases in nulls or small differences in reconciliations — send to the data team with low/normal priority for scheduled investigation.
  • Informational alerts: detected drift trends without immediate impact — log for weekly review and include in the data quality meeting.

Other useful tactics: silence windows during planned operations (deploys, reloads), grouping alerts by root cause (alerts from the same table/partition aggregated into a single incident) and dynamic thresholds that adjust tolerance according to historical variability. Automate simple playbooks — for example, restart a job, reprocess a specific partition, or trigger an automatic rerun up to 3 attempts — so that only problems requiring manual intervention generate human alerts.

Operation, governance and BI integration

Observability is technical, but it requires procedural discipline. Define SLAs for freshness and accuracy per dataset (e.g.: online sales table — freshness ≤ 15 minutes, completeness ≥ 99%). Version schemas and change policies that require automated tests before promoting changes to production.

Document runbooks with clear steps: check pipeline logs, consult the health table, restore a partition from backups or reprocess the source, notify stakeholders. For each runbook, define a maximum time for each step and an escalation contact. Record post‑incident metrics (MTTD, MTTR) and analyze quarterly to reduce recurrence.

Integrate these metrics into Power BI reports on data trust. A 'Data Health' pane for managers should show percentages of SLA compliance, ongoing incidents and quality trends by domain. Make the information actionable: each row of the pane should have a direct link to the corresponding notebook or runbook page. This turns observability into a visible signal for business teams, preventing decisions made with data that do not meet minimum requirements.

In summary

  • Prioritize simple, actionable metrics: freshness, completeness, integrity and schema drift.
  • Implement light checks per ingestion and scheduled heavy checks; record everything in a health table in the Lakehouse.
  • Practical architecture: instrumentation in pipelines, metrics storage in OneLake/Log Analytics and Power BI dashboards with integrated alerts.
  • Define clear SLAs, runbooks and alert categorization to avoid false positives and alert fatigue.
  • Measure impact: reduce resolution time and regain users’ trust with stable, predictable data.

Implementing data observability in Fabric is an incremental effort: start with critical assets, prove value by reducing incidents and expand. If your organization still treats data errors as 'inevitable accidents', propose a 6‑week pilot on a critical pipeline — with pre and post metrics to demonstrate ROI. In many cases, a simple pilot covering 3 tables and 2 pipelines reveals measurable improvements in less than a month.

Which data are critical in your business that, if they failed tomorrow, would stop operations? Let’s discuss how to design an observability pilot that delivers impact in 6 weeks.

← Back to insights
Let's talk?

Ready to transform your data?

Book a free 30-minute meeting and find out how we can help your team make better decisions.

Book a Free Meeting
bConcepts