How to build a quantile regression model in Machine Learning
Quantile regression allows predicting percentiles of the conditional distribution of the target variable, rather than just the mean. This is especially useful when the dispersion of the variable depends on X (heteroscedasticity) or when you want to assess risks and conditional confidence intervals — for example, estimating minimum, median and maximum costs for an operation. In this example we show how to train several quantile regression models in Python using scikit-learn, because it provides a straightforward and efficient solution to obtain quantiles such as 0.1, 0.5 and 0.9, and why this improves decisions in scenarios with uncertainty.
Prerequisites
- Python 3.8+ with pip installed
- Libraries: scikit-learn, numpy, pandas, matplotlib
- Basic knowledge of linear regression, train/test splitting and data handling with pandas
Step 1: Why use quantile regression
Quantile regression estimates a quantile q (for example 0.1, 0.5, 0.9) of the conditional distribution of y given X. Unlike least squares regression, which minimizes squared error and provides the conditional expectation E[y|X], quantile regression models the tails and the median directly. This is valuable when decisions depend on variability — for example, a risk manager may want to know the 0.95 quantile of potential losses; a budget planner may compare estimated costs for the 0.1 and 0.9 quantiles to prepare contingencies.
Concrete example: if the variance of y increases with X, mean regression tends to "stay in the middle" and underestimates the extremes. Quantile regression shows how intervals widen with X and allows measuring empirical coverage (the proportion of observations that fall between two quantiles). For quantiles 0.1–0.9 the theoretical target for coverage is 0.8 (80%).
Step 2: Install dependencies
Install the required libraries. In the terminal, run:
pip install scikit-learn numpy pandas matplotlib
scikit-learn includes QuantileRegressor since version 1.0; check the version with pip show scikit-learn if necessary. On large datasets you may prefer different solvers and regularization options to optimize performance.
Step 3: Prepare example data
To illustrate heteroscedasticity we create a synthetic dataset with 500 observations; the amplitude of the noise grows with X. This makes it visible why quantiles diverge from the mean.
import numpy as np
import pandas as pd
from sklearn.model_selection import train_test_split
np.random.seed(0)
X = np.linspace(0, 10, 500)
noise = np.random.randn(500) * (0.5 + 0.5 * X)
y = 2.0 + 0.5 * X + noise
df = pd.DataFrame({'X': X, 'y': y})
X = df[['X']]
y = df['y']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
The training sample will have 400 points and the test sample 100 points. This is a sufficient size to demonstrate quantile behavior without requiring much computing power.
Step 4: Train quantile regression with scikit-learn
scikit-learn offers QuantileRegressor (since v1.0). We train separate models for the 0.1, 0.5 and 0.9 quantiles. The alpha parameter controls L2 regularization: values like 0.0–1.0 are common; 0.0 means no regularization. The 'highs' solver is stable and fast for linear problems.
from sklearn.linear_model import QuantileRegressor
quantiles = [0.1, 0.5, 0.9]
models = {}
for q in quantiles:
qr = QuantileRegressor(quantile=q, alpha=0.0, solver='highs')
qr.fit(X_train, y_train)
models[q] = qr
Note that a more robust workflow includes feature normalization and validation of alpha via cross-validation if you have multiple features.
Step 5: Make predictions and build the conditional interval
Use the models to predict the quantiles on the test set; the difference between the 0.9 and 0.1 quantiles approximates a robust conditional interval. This provides estimates of the interval width dependent on X.
import numpy as np
preds = {q: models[q].predict(X_test) for q in quantiles}
median_pred = preds[0.5]
lower_pred = preds[0.1]
upper_pred = preds[0.9]
For example, for X=10 the (0.1–0.9) interval may be significantly wider than for X=0; this reflects increasing uncertainty with X.
Step 6: Visualize results (example)
A visualization helps verify whether the quantiles capture heteroscedasticity. A typical plot shows observed points and the three quantile curves. If the model is correct, points should be dense within the band (0.1–0.9) and the band should widen with X.
import matplotlib.pyplot as plt
plt.scatter(X_test, y_test, s=10, alpha=0.6, label='observado')
# Ordenar para linha
order = np.argsort(X_test.values.ravel())
xx = X_test.values.ravel()[order]
plt.plot(xx, lower_pred[order], color='red', label='quantil 0.1')
plt.plot(xx, median_pred[order], color='green', label='quantil 0.5')
plt.plot(xx, upper_pred[order], color='red', linestyle='--', label='quantil 0.9')
plt.legend()
plt.xlabel('X')
plt.ylabel('y')
plt.title('Regressão quantílica: intervalos condicionais')
plt.show()
Step 7: Evaluate quantile performance
To evaluate we use the quantile loss (pinball loss) and the empirical coverage of the interval. The pinball loss is asymmetric and penalizes deviations appropriately for each side of the quantile.
def pinball_loss(y_true, y_pred, q):
err = y_true - y_pred
return np.mean(np.maximum(q * err, (q - 1) * err))
for q in quantiles:
loss = pinball_loss(y_test.values, preds[q], q)
print(f'Quantile {q}: pinball loss = {loss:.4f}')
coverage = np.mean((y_test.values >= lower_pred) & (y_test.values <= upper_pred))
print(f'Cobertura entre 0.1 e 0.9: {coverage:.3f}')
As a guideline, for our synthetic example the empirical coverage should be near 0.8 (for example between 0.75 and 0.85). Values far outside this range indicate under- or overfitting or modeling issues.
Check the outcome
Confirm that the plots show increasing dispersion with X and that the empirical coverage of the 0.1–0.9 interval is close to 0.8. Avoid common mistakes: using a single mean regression when you need intervals, not validating alpha (over-regularization can flatten quantiles), or forgetting to transform/normalize features when you have multiple columns with different scales.
Conclusion
Quantile regression in Machine Learning allows estimating conditional intervals and managing uncertainty in a practical and interpretable way. Recommended next steps: experiment with multiple features, choose alpha via cross-validation, test non-linear models (e.g., GradientBoostingQuantile or models based on quantile loss) and validate quantile calibration on real data. Practical tip: always validate empirical coverage to ensure quantiles are well calibrated before using predictions in operational decisions.