(+351) 21 24 10006  ·  info@bconcepts.pt
Carnaxide, Lisbon
CDC in Microsoft Fabric — robust, practical pipelines
Data Engineering

CDC in Microsoft Fabric — robust, practical pipelines

João Barros 08/10/2026 9 min

Data mutates every second — the value is in capturing that mutation reliably and actionably.

Why CDC matters in Lakehouse architectures

Change Data Capture (CDC) has moved from a backstage technique to a central piece of modern analytics platforms. In an era where operational decisions and strategic reporting require fresh data, moving from daily batch loads to near-real-time change capture transforms processes: fraud detection, inventory updates, user experience metrics and financial reporting with reduced latency.

CDC in Microsoft Fabric — robust, practical pipelines

In the context of Microsoft Fabric, where the Lakehouse and OneLake act as the single source of truth for BI and AI, integrating CDC correctly ensures that data models in Power BI and Machine Learning pipelines have consistent inputs. The challenge is not just moving events; it is preserving order, ensuring idempotency, managing schema evolution and minimizing operational costs — all without sacrificing latency and integrity SLAs.

Beyond the obvious benefits of fresh data, a well-designed CDC pipeline reduces indirect costs: fewer reworks in reports, fewer manual reconciliation queries and business decisions made with confidence. In medium-sized organizations, it is often possible to reduce decision cycles from hours to minutes and cut 10–30% of the human effort associated with data corrections.

Ingestion choices: streaming vs micro-batch in Fabric

The first practical decision is the ingestion pattern. Micro-batch (for example, incremental queries via Azure Data Factory/Power Query on a 1–5 minute cadence) is simple to implement and sufficient for many analytical scenarios with tolerant latency. In contrast, streaming (Debezium/Kafka/Event Hubs → Spark Structured Streaming in Fabric) is appropriate when sub-minute latency and per-event processing are required — for example, payment systems, real-time anomaly detection or instant inventory updates during campaigns.

In Fabric, a common pattern is: a CDC-capable source (SQL Server, PostgreSQL, Cosmos DB), a CDC producer (Debezium or native protocol), transport (Apache Kafka or Azure Event Hubs) and a Spark Structured Streaming consumer writing to Delta Lake in the Lakehouse. For simpler scenarios, micro-batch pipelines with watermarking and incremental queries reduce complexity and costs while still achieving latencies on the order of 1–5 minutes.

Some practical rules to decide: if the business requires average latency below 1 minute and predictable peaks up to tens of thousands of events per second, choose streaming. If acceptable latency is 1–15 minutes and the team favors operational simplicity, micro-batch is often the right choice. There is also a middle ground: 30–60 second micro-batches that offer a good balance between latency and operational cost.

Idempotency, ordering and conflict resolution

In a world of retries, duplicates and variable latency, designing idempotent pipelines is mandatory. Practical strategies include using natural keys or surrogate keys, a sequence/LSN (Log Sequence Number) field and a change timestamp. When writing to Delta tables, the pattern "merge by key using sequence" allows applying only the most recent change, discarding duplicates or reorderings.

A common CDC event schema contains: pk, op_type (I/U/D), change_ts, lsn, payload and possibly an is_tombstone field. The consumer performs a MERGE by pk with an update condition only when lsn or change_ts is greater than what is currently persisted. Logical example of the condition: WHEN MATCHED AND incoming.lsn > target.lsn THEN UPDATE SET ...

To handle reorderings, keep tolerance windows (for example, 15 minutes) during which late events are accepted to update records, and records arriving outside that window go to a "late arrivals" table for manual or automatic reconciliation. In systems with zero tolerance for temporal inconsistencies — for example, inventory counts synchronized with POS — consider additional mechanisms like producer-side read confirmations or transaction-based versions.

Delete handling deserves attention: do not assume that a delete in the source means immediately deleting the record in the Lakehouse. Instead, use tombstones (is_deleted=TRUE, change_ts) and retention/purge policies, which allow auditing and reconciling before performing a final VACUUM.

Managing schema evolution in CDC pipelines

Schema evolution is the main source of breakages in CDC pipelines. Adding columns, changing types or renaming fields can break upstream consumers and downstream reports. There are two broad tactics: permissive and explicit. The permissive approach uses self-describing formats (AVRO/JSON with schema registry) and write operations that support schema-on-read/merge; the explicit approach introduces schema versions and clear contracts between product owners and data teams.

Concrete practices that reduce risk: 1) enforce schema compatibility in the registry (backwards/forwards/fully compatible as appropriate); 2) use new fields as optional and with default values; 3) avoid direct renames — instead, introduce the new column alongside the old one and plan removal after a coexistence period (for example, 90 days).

In Fabric, using a schema registry (Confluent or Azure Schema Registry) and ensuring Delta writes involve controlled schema merging is pragmatic. For testing, automate scenarios in staging: change column type integer->bigint, add arrays/structs, and validate consumers. Always include backfills when a column becomes non-nullable — measuring the cost of backfill in I/O and time helps decide the migration window.

Compaction, small files and cost management

A common denominator in CDC pipelines is the appearance of many small files, which degrades both query performance and increases I/O costs. The strategy is to write in controlled batches and/or enable coalescing and periodic compaction techniques. In the Delta context, running OPTIMIZE/compact operations (or using Fabric's optimise mechanism when applicable) reduces files and improves read time.

A recommended file threshold: ideally Parquet files between 128 MB and 512 MB for analytical queries. For streaming workloads, group events by time intervals (e.g., 1–5 minutes) and force nightly compaction to balance latency and cost. For example, writing in 2-minute micro-batches with coalesce into ~256 MB files and running deep compaction once a day typically reduces read latency by 30–70%.

Do not forget retention policies and controlled VACUUM to clean obsolete files without compromising pending transactions. In environments that retain versions for compliance, plan additional space: keeping 7–14 days of versions can increase storage by 10–30%, depending on change rate.

Observability, testing and SLAs for CDC pipelines

Without metrics and automated tests, a CDC pipeline is a black box. Instrument each stage: event counters read, end-to-end latency (P50/P95/P99), error rate, duplicate counters and file size/count. Structured logs and metrics exported to a monitoring system (Application Insights, Log Analytics or Grafana) enable alerts for delay thresholds (e.g., lag > 5 min) or continuous error.

Concrete metrics to track (examples): average and peak throughput (events/sec), P95 lag < 5 minutes, duplicate rate < 0.05%, daily reconciliation divergence < 0.02%. Active alerts should include operational notification when lag exceeds SLAs, when the number of small files rises 3x in 24 hours, or when the error rate per minute exceeds a threshold (e.g., > 10/min).

Health checks should include contract tests, schema regression tests and checksum comparisons between source and destination. Automate tests with CI pipelines that validate merges, streaming runs in sandbox and failure scenarios — for example: consumer restart, event duplication, and schema change in staging. Reconciliation procedures can be simple: compare hourly counts and sums of critical fields (e.g., total_sales) and recompute checksums by pk; discrepancy tolerance can be defined according to data criticality.

Mini practical case: implementing CDC in an 80-employee company

In an 80-person company with an e-commerce platform, the transactional database contains 10 million customer records and grows by 100k events/day. The goal: reduce inventory and sales dashboard update latency from 24 hours to under 5 minutes while keeping operational cost moderate.

Implemented solution: enable CDC on SQL Server (source), use Debezium to extract changes to Azure Event Hubs, and a Spark Structured Streaming job in Fabric to consume events and write to Delta tables in the Lakehouse. They also implemented a schema registry and metrics in Log Analytics. Details and results after 3 months:

  • Average event rate: 2.5k events/sec; peak 7k/s during promotions.
  • Average end-to-end latency: 3.8 minutes (objective <5 min achieved). P95: 7.2 minutes during peaks.
  • Operational report refresh in Power BI reduced from 45 minutes to 6 minutes.
  • Incremental compute costs: ~18% monthly cost increase, mainly due to clusters reserved for streaming; this cost was offset by a 22% reduction in stock loss and a 4% improvement in conversion during campaigns.
  • Operation: nightly compaction reduced small files by 85% and decreased I/O by 40%; daily reconciliation showed 99.98% agreement between source and Lakehouse.

This mini-case illustrates that, even in small teams, with pragmatic choices (native CDC, Debezium, Event Hubs and Spark in Fabric) you can achieve a qualitative leap in business response time without exponential cost increases. The key was investing in the first 2–4 sprints in observability and automated tests — this reduced production incidents by 70% in the second quarter.

A healthy CDC pipeline is not the one that transmits the most events per second — it is the one that delivers correct data, at the right time and in a predictable way.

Practical checklist to implement CDC in Microsoft Fabric

Before starting, confirm these points with your technical and product teams:

  • Source with native CDC or adapters (Debezium/CDC connector) and stable key schema.
  • Transport decision: Kafka/Event Hubs for streaming, or micro-batch pipelines for incremental ingestion.
  • Message design with sequence/LSN field and change timestamp; include is_tombstone for deletes.
  • Idempotent write mechanism (Delta MERGE by key + sequence condition).
  • Schema evolution policies (schema registry or controlled versioning) and automated tests.
  • Compaction routines and retention policies defined to minimize small files and costs.
  • Metrics and alerts configured: lag, throughput, errors, files, and daily reconciliation.

In summary

  • Choose streaming when you need sub-minute latency; opt for micro-batch when latency allows simplicity and lower cost.
  • Design idempotency with key + sequence/LSN; use conditional merges to avoid regressions from reordering.
  • Plan schema evolution with a schema registry and controlled migrations; avoid direct renames.
  • Mitigate small files with compaction and file thresholds; optimize for Parquet files of 128–512 MB.
  • Automate observability, testing and reconciliation to maintain reliable operational SLAs.

Conclusion and next steps

Implementing CDC in Microsoft Fabric is a combination of architectural best practices and operational discipline. Decisions — transport, ingestion strategy, schema management and compaction policies — directly influence the quality of data consumed by Power BI and AI models. For teams starting out, the pragmatic path is to prove a minimal viable flow: enable CDC at the source, move events to a topic in Event Hubs, and validate a simple micro-batch or streaming pipeline that writes to Delta with idempotent merge operations.

Recommended next steps are: 1) create a staging prototype with a subset of critical data; 2) implement basic metrics and alerts; 3) automate contract and reconciliation tests; 4) validate costs with representative workloads. These steps reduce risk and allow rapid iteration on a robust platform.

Would you like to discuss a concrete case for your organization or validate an architectural design for your CDC pipeline in Microsoft Fabric?

← Back to insights
Let's talk?

Ready to transform your data?

Book a free 30-minute meeting and find out how we can help your team make better decisions.

Book a Free Meeting
bConcepts