(+351) 21 24 10006  ·  info@bconcepts.pt
Carnaxide, Lisbon

How to generate stratified samples in Python for data: step by step

João Barros 26 de September de 2026 4 min read

Generating stratified samples in Python for data helps ensure that the proportions of relevant classes or categories are preserved in the sample, making analyses and models more representative. This guide shows how to create simple and practical stratified samples, for example for train/test or validation, and explains common mistakes to avoid.

Prerequisites

  • Python 3.8+ installed
  • Libraries: pandas and scikit-learn (sklearn)
  • CSV file or DataFrame with at least one categorical column to stratify by

Step 1: Why use stratified sampling?

Stratified sampling ensures that the distribution of categories (for example, class labels) in the sample reflects the distribution in the population. It is useful when there are imbalanced classes and when we want to avoid bias in model evaluation or data exploration.

Step 2: Install and import libraries

Install pandas and scikit-learn if they are not already available. Then import what is needed. Avoid installing on every run; use a virtual environment for reproducibility.

pip install pandas scikit-learn
import pandas as pd
from sklearn.model_selection import train_test_split

Step 3: Load data and identify the stratum

Read the CSV into a DataFrame and choose the column that defines the strata (for example, 'target' or 'categoria'). Check the original distribution before sampling.

# Minimal example
df = pd.read_csv('dados.csv')  # ou use um DataFrame já carregado
print(df['target'].value_counts(normalize=True))

Step 4: Stratified sample for train/test with sklearn

Use train_test_split with the stratify argument to maintain proportions. Set test_size or train_size as needed. This method is the simplest and most robust for an initial split.

# Separate X and y if needed
X = df.drop(columns=['target'])
y = df['target']

# Stratified split: 80% train, 20% test
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y
)

# Recreate full DataFrames if useful
train_df = X_train.copy()
train_df['target'] = y_train

test_df = X_test.copy()
test_df['target'] = y_test

Step 5: Stratified sampling by group or multiple columns

If the stratum is the combination of multiple columns (for example, 'sexo' + 'idade_cat') first create a composite column and then stratify by it. When stratifying by large groups, pay attention to the minimum number of samples per stratum.

# Create a composite stratum
df['estrato'] = df['sexo'].astype(str) + '_' + df['idade_cat'].astype(str)

# Check minimum sizes per stratum
print(df['estrato'].value_counts().head())

# Use stratify with the composite column
X = df.drop(columns=['target'])
y = df['target']
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.25, random_state=1, stratify=df['estrato']
)

Step 6: Sample a stratified fraction (e.g., 10% of each stratum)

If you want a sample that contains 10% of each stratum, use groupby + sample in pandas. This allows controlling the fraction per stratum, but be careful with strata that have few observations.

# Sample 10% of each stratum (with replace=False)
sample_frac = 0.10
sampled = df.groupby('estrato', group_keys=False).apply(
    lambda x: x.sample(frac=sample_frac, random_state=42)
).reset_index(drop=True)

# If there are very small strata you can use replace=True or filter

Step 7: Common mistakes and how to avoid them

Frequent mistakes include stratify with strata that have only 1 observation (train_test_split fails) and forgetting to keep the same random_state for reproducibility. Always check value_counts before and after sampling.

# Check distributions
print('Original:', df['target'].value_counts(normalize=True))
print('Train:', train_df['target'].value_counts(normalize=True))
print('Test:', test_df['target'].value_counts(normalize=True))

Verify the result

Confirm that the proportions per stratum in the sample match those of the population. Compare value_counts(normalize=True) between original, train and test or between original and the stratified sample.

def comparar_proporcoes(col, *dfs):
    for i, d in enumerate(dfs, 1):
        print(f'Dataset {i}:')
        print(d[col].value_counts(normalize=True))
        print()

comparar_proporcoes('target', df, train_df, test_df)

Conclusion

Stratified sampling in Python with pandas and sklearn ensures representativeness by categories/strata and reduces bias in analyses and model validation. Next steps: try stratified k-fold with sklearn.model_selection.StratifiedKFold and handle very small strata (aggregate or use oversampling). Tip: always confirm distributions before and after sampling.