AI-901: implement data preprocessing pipelines in Microsoft Foundry
I will explain how to implement data preprocessing pipelines in Microsoft Foundry — a practical skill useful for the AI-901. Correctly preparing data before inferring or training models is critical both for result accuracy and for safe operationalization in production.
What you need to know
Data preprocessing means transforming raw data into a form suitable for AI models or inference flows. In the context of Microsoft Foundry (part of the Fabric/OneLake ecosystem), this includes: cleaning (removing duplicates, handling nulls), normalization/scaling, tokenization for text, resizing/normalization for images, and transforming categorical columns into numerical representations.
Example: you have transaction records with free-text fields, dates in multiple formats, and an amount column with a currency symbol. A typical preprocessing pipeline will: (1) unify date formats, (2) remove symbols and convert amounts to float, (3) impute or remove null values, (4) extract temporal features (day of week, hour), (5) export a normalized file/table for consumption by a model or an inference endpoint.
How it works
In Foundry you implement preprocessing pipelines by assembling components that transform data in reproducible steps. You typically use:
- Datasets or tables in OneLake as source/target.
- Transforms/activities from Foundry to run code (Python, PySpark) or visual operators.
- Orchestration tools to schedule and version pipelines.
Logical flow of a pipeline:
- Ingestion: load raw data into a staging table.
- Validation: automatic checks (schema, types, ranges).
- Cleaning: remove duplicates, normalize texts and numbers.
- Transformation: create features (one-hot, simple embeddings, normalization).
- Export: write table ready for training/inference and record metadata.
In practice
Here is a simplified example in pseudo-Python/PySpark that you can adapt as a transform in Foundry. It assumes you read a OneLake table called transactions_raw and will write transactions_clean.
# Exemplo simplificado de transform (pseudo-PySpark)
from pyspark.sql import functions as F
df = spark.read.table('oneLake.transactions_raw')
# Normalizar datas
df = df.withColumn('transaction_date', F.to_timestamp(F.col('date_str'), 'yyyy-MM-dd HH:mm:ss'))
# Limpar montantes: remover símbolos e converter
df = df.withColumn('amount_clean', F.regexp_replace(F.col('amount'), '[^0-9.-]', '').cast('double'))
# Imputar valores nulos (ex.: 0 para montantes, 'unknown' para categorias)
df = df.fillna({'amount_clean': 0.0, 'category': 'unknown'})
# Extrair features temporais
df = df.withColumn('day_of_week', F.dayofweek('transaction_date'))
# One-hot simples para categoria (exemplo conceptual)
categories = ['food','travel','utilities']
for c in categories:
df = df.withColumn(f'cat_{c}', F.when(F.col('category') == c, 1).otherwise(0))
# Gravar tabela limpa
df.write.mode('overwrite').saveAsTable('oneLake.transactions_clean')
Practical notes:
- Validate the schema before and after; Foundry allows recording schema and lineage for auditing.
- If you work with text for LLMs, add specific tokenization/cleaning and limit size to control inference costs.
- For images, include resizing, format conversion and normalization of pixel values.
Common mistakes
- Not versioning transforms: changing code without versioning compromises reproducibility. Use the built-in version control in Foundry to record each transform.
- Applying different transformations in training and inference: if the training pipeline normalizes or encodes data in way X, the inference pipeline must apply exactly the same logic and parameters.
- Ignoring input validation: assuming data always arrives in the expected format (e.g., dates or symbols) leads to production failures. Implement checks and alerts.
How to practice
To study this skill, use Microsoft’s official free resources:
- Microsoft OFFICIAL Practice Assessment (free) — it’s your reference to get familiar with the exam format and measured areas.
- Official AI-901 study guide (free) — contains the skills measured and recommended resources.
In practice, create a small project in Foundry (or a Fabric/OneLake environment you have access to): ingest a CSV, implement a transform in Python/PySpark with validations and export the cleaned table. Test the same pipeline with example data and with "noisy" data to verify robustness.
In summary
- Preprocessing prepares raw data for training and inference: cleaning, normalization, feature engineering and export.
- In Microsoft Foundry you implement pipelines as reproducible transforms, connected to OneLake and versioned for auditability.
- Avoid discrepancies between training and inference by applying exactly the same transformations; validate inputs and version code.
- Practice with the official Practice Assessment and the Microsoft study guide and implement real labs in Foundry to consolidate.