DP-700: how to implement data retention and lifecycle policies in Fabric
I will teach you the skill of implementing data retention and lifecycle policies in Microsoft Fabric, an ability required in DP-700 that ensures compliance, reduces storage costs and maintains performance. In practice, this means configuring rules that move, archive, or delete data based on its age, sensitivity and usage.
What you need to know
Retention and lifecycle policies define how data is managed from creation to deletion. In the context of Fabric and a Data Lakehouse (OneLake/Delta), this includes:
- Retention: how long to keep data for legal, audit or business reasons.
- Tiering/Archiving: move data to cheaper storage tiers (for example, archive older files) or to formats optimized for infrequent querying.
- Secure deletion: remove data irreversibly when the retention period expires, respecting legal and security retention policies.
Practical example: an application logs dataset can be kept for 90 days in hot storage for fast analysis, then moved to archive for another 2 years and finally deleted after 2 years and 90 days.
How it works (practical step-by-step)
Follow a practical process, with concrete actions you apply in Fabric and associated artifacts (OneLake, Delta tables, Dataflows, Synapse).
-
Inventory the data and define requirements: categorize datasets by sensitivity, access frequency and legal requirements. Use the data catalog to record metadata: owner, classification and retention period.
-
Choose a storage strategy: decide when to use tiers (hot/warm/cold), Delta files optimized for read vs files for archiving (compacted parquet). For transactional data, keep Delta versioning; for logs rarely accessed, consider copies in formats optimized for archiving.
-
Implement automatic ageing rules: create pipelines (Dataflow, Synapse Spark or Notebooks in Fabric) that perform date-based operations:
# Exemplo conceptual em pseudocódigo PySpark from delta.tables import DeltaTable # carregar a tabela Delta delta = DeltaTable.forPath(spark, "/lakehouse/areas/prod/logs") # identificar partições antigas old_df = spark.read.format("delta").load("/lakehouse/areas/prod/logs") \ .filter("event_date < date_sub(current_date(), 90)") # mover dados para arquivamento (escrever em local de arquivamento) old_df.write.mode("append").parquet("/archive/logs/year=2023") # apagar das tabelas activas (gerir transacção Delta) delta.delete("event_date < date_sub(current_date(), 90)")This script illustrates: identify old data, copy to archive and delete from the active table ensuring consistency with Delta.
-
Automate and schedule tasks: use pipelines/Jobs in Fabric to run these actions at regular intervals. Add logs and alerts for failures.
-
Record and monitor changes: keep an inventory with timestamps of movements and deletions. Use the catalog and audits (audit logs) for compliance evidence.
-
Implement secure deletion: when a record reaches end of life, ensure it is deleted from all locations (OneLake, Delta snapshots, backups). In Delta, manage snapshots and vacuum carefully:
# Exemplo conceptual: executar VACUUM com retenção segura # vacuum('/lakehouse/areas/prod/logs', retentionHours=168) # 7 diasDo not reduce VACUUM retention below legal needs; this can remove snapshots that other operations still reference.
Common mistakes
- Deleting snapshots or files too early: reducing the VACUUM retention window to reclaim space without assessing dependencies can break reproducibility and recovery.
- Not synchronizing catalog and storage: moving files without updating the catalog causes failed queries and confusion about data location.
- Failing to fully delete: forgetting backups, copies or logs that contain sensitive data prevents meeting deletion requirements and can create legal risks.
How to practice
Practice with a simple lab: create a Delta table with simulated log data, implement a notebook that archives and deletes old records, and use VACUUM and the catalog to verify results. For formal preparation, take the official free Microsoft Practice Assessment and consult the official DP-700 study guide (both free). These resources help you assess knowledge gaps without resorting to unauthorized material.
In summary
- Retention and lifecycle policies protect compliance, reduce costs and maintain analytics environment performance.
- Implement a process: inventory → rule design → automation → monitoring → secure deletion.
- Be cautious with VACUUM and Delta snapshots; always synchronize catalog and storage.
- Use Microsoft’s official resources (Practice Assessment and free study guide) to validate and strengthen practice.