How to do customer clustering with Azure ML: step by step
This tutorial shows how to do customer clustering in Azure AI & Machine Learning for marketing segmentation and exploratory analysis. I explain why segmentation helps personalize campaigns and how to build a simple pipeline that pre-processes data, trains a clustering model (KMeans) and publishes an endpoint for inference.
Prerequisites
- Azure account with permissions to create resources (Azure Machine Learning workspace).
- Azure CLI and az ml extension installed or access to Azure Machine Learning Studio.
- Python 3.8+ and packages: azure-ai-ml, pandas, scikit-learn, joblib.
- Customer dataset (CSV) with simple numeric and categorical columns.
Step 1: Create the workspace and configure the environment
We need an Azure Machine Learning workspace and a Python environment with dependencies. If you already have the workspace, proceed. Otherwise, create it with the CLI and configure credentials locally.
# Login e seleção de subscrição
az login
az account set --subscription ""
# Criar resource group e workspace (se necessário)
az group create -n rg-ml -l westeurope
az ml workspace create -w ml-workspace -g rg-ml -l westeurope
# Criar ambiente virtual (opcional)
python -m venv .venv
source .venv/bin/activate # ou .venv\Scripts\activate no Windows
pip install azure-ai-ml pandas scikit-learn joblib
Step 2: Prepare and load the data
Load the customers CSV, handle missing values and convert categorical variables. For clustering, normalizing numeric attributes is essential to avoid bias.
import pandas as pd
from sklearn.preprocessing import StandardScaler
from sklearn.preprocessing import OneHotEncoder
# Carregar dados
df = pd.read_csv('clientes.csv')
# Selecionar colunas exemplo
num_cols = ['idade','rendimento','gastos_mensais']
cat_cols = ['regiao']
# Preencher NA e transformar
df[num_cols] = df[num_cols].fillna(df[num_cols].median())
df[cat_cols] = df[cat_cols].fillna('Desconhecido')
# One-hot para categóricas
ohe = OneHotEncoder(sparse=False, handle_unknown='ignore')
cat_enc = ohe.fit_transform(df[cat_cols])
cat_df = pd.DataFrame(cat_enc, columns=ohe.get_feature_names_out(cat_cols))
# Normalizar numéricos
scaler = StandardScaler()
num_scaled = scaler.fit_transform(df[num_cols])
num_df = pd.DataFrame(num_scaled, columns=num_cols)
# Concat
X = pd.concat([num_df.reset_index(drop=True), cat_df.reset_index(drop=True)], axis=1)
X.to_csv('clientes_prepared.csv', index=False)
Step 3: Train the KMeans model locally
Training locally helps iterate quickly. KMeans is simple and interpretable for segmentation. Try different values of k and use the elbow method to choose.
from sklearn.cluster import KMeans
import matplotlib.pyplot as plt
# Carregar preparado
X = pd.read_csv('clientes_prepared.csv')
# Método do cotovelo
inertia = []
K = range(2,8)
for k in K:
kmeans = KMeans(n_clusters=k, random_state=42)
kmeans.fit(X)
inertia.append(kmeans.inertia_)
# Gravar e inspecionar (pode visualizar o gráfico)
import joblib
k_opt = 4 # escolha após inspeção do cotovelo
model = KMeans(n_clusters=k_opt, random_state=42)
model.fit(X)
joblib.dump((model, scaler, ohe), 'kmeans_clientes.pkl')
Step 4: Create a scoring script and package for deployment
To publish an endpoint in Azure ML, create a script that loads the model and accepts new records in JSON/CSV to predict the cluster. Keep the code minimal and robust against common errors (missing columns).
# scoring.py
import json
import pandas as pd
import joblib
model, scaler, ohe = joblib.load('kmeans_clientes.pkl')
def preprocess(input_df):
# Supondo mesma lógica do treino
num_cols = ['idade','rendimento','gastos_mensais']
cat_cols = ['regiao']
input_df[num_cols] = input_df[num_cols].fillna(input_df[num_cols].median())
input_df[cat_cols] = input_df[cat_cols].fillna('Desconhecido')
cat_enc = ohe.transform(input_df[cat_cols])
cat_df = pd.DataFrame(cat_enc, columns=ohe.get_feature_names_out(cat_cols))
num_scaled = scaler.transform(input_df[num_cols])
num_df = pd.DataFrame(num_scaled, columns=num_cols)
X = pd.concat([num_df.reset_index(drop=True), cat_df.reset_index(drop=True)], axis=1)
return X
def predict(json_input):
df = pd.DataFrame(json.loads(json_input))
X = preprocess(df)
clusters = model.predict(X)
return json.dumps({'clusters': clusters.tolist()})
Step 5: Publish the model as an endpoint in Azure Machine Learning
Use the azure-ai-ml SDK to register the model and create an endpoint for real-time inference. Below are the essential steps; adjust resources and policies as needed.
from azure.ai.ml import MLClient
from azure.identity import DefaultAzureCredential
from azure.ai.ml.entities import Model, ManagedOnlineEndpoint, ManagedOnlineDeployment
subscription_id = ''
resource_group = 'rg-ml'
workspace = 'ml-workspace'
mlclient = MLClient(DefaultAzureCredential(), subscription_id, resource_group, workspace)
# Registar o ficheiro do modelo
model = mlclient.models.create_or_update(Model(name='kmeans-clientes', path='kmeans_clientes.pkl'))
# Criar endpoint gerido
endpoint = ManagedOnlineEndpoint(name='endpoint-kmeans', auth_mode='key')
mlclient.begin_create_or_update(endpoint).result()
# Criar deployment com imagem que tenha Python e dependências (exemplo mínimo)
deployment = ManagedOnlineDeployment(
name='deploy-v1',
endpoint_name=endpoint.name,
model=model,
environment={'conda_file': 'env.yml'},
code_configuration={'code': '.', 'scoring_script': 'scoring.py'}
)
mlclient.begin_create_or_update(deployment).result()
# Traffic split
endpoint.traffic = {'deploy-v1': 100}
mlclient.begin_create_or_update(endpoint).result()
Verify the result
Test the endpoint with a sample payload and confirm you receive valid clusters. Also check usage metrics and logs in Azure Machine Learning Studio to troubleshoot common errors (serialization errors, missing columns).
import requests
# Obter chave e URL do endpoint via MLClient ou portal
endpoint_url = 'https:///score'
api_key = ''
headers = {'Content-Type': 'application/json', 'Authorization': f'Bearer {api_key}'}
payload = [
{'idade': 34, 'rendimento': 25000, 'gastos_mensais': 300, 'regiao': 'Norte'}
]
resp = requests.post(endpoint_url, headers=headers, json=payload)
print(resp.text) # deve mostrar clusters
Conclusion
You performed customer clustering in Azure AI & Machine Learning: you prepared data, trained KMeans, wrote a scoring script and published an endpoint. Next steps: try other algorithms (DBSCAN, GaussianMixture), evaluate cluster stability and integrate segmentation with campaigns. Tip: always test with real cases and verify that the variables used are relevant for marketing.