How to implement percentile-based latency alerts in Real-Time Analytics
This tutorial shows how to implement alerts that monitor latency percentiles in Real-Time Analytics — useful to detect performance degradations per user or per service. We will compute percentiles over sliding windows and emit alert events when the value exceeds a threshold.
Prerequisites
- Account and instance with a streaming system (e.g.: Apache Kafka, Event Hubs) and a real-time processing engine (e.g.: Azure Stream Analytics, Flink, Spark Structured Streaming).
- Event source with timestamp, request_id, service, latency_ms.
- Notification tool (e.g.: webhook, Azure Function, Prometheus Alertmanager) to receive alerts.
Step 1: Choose windowing and percentile strategy
Decide the time window (e.g.: 1 minute, sliding every 15s) and the percentile to monitor (e.g.: 95th). A sliding window reduces false positives and gives quick visibility. Example: 1-minute window with 15s slide to compute the 95th percentile of latency_ms per service.
Step 2: Normalize and validate events
Ensure each event has a correct timestamp and latency_ms is numeric; remove absurd outliers (e.g.: negative or > 1 hour). This prevents incorrect percentile calculations.
// Pseudocode for validation (adapt to your engine) // receives event {timestamp, request_id, service, latency_ms} if (!timestamp || !service || latency_ms == null) drop(event) if (latency_ms < 0 || latency_ms > 3600000) drop(event) // more than 1h invalid emit(validated_event)
Step 3: Compute the percentile in streaming
Use the engine's percentile function (e.g.: APPROX_PERCENTILE, percentile_approx, approxQuantile). In the example below I use pseudo-SQL compatible with engines that support sliding window aggregations.
// Conceptual SQL example for a streaming engine that supports sliding windows SELECT
TUMBLE_START(event_time, INTERVAL '15' SECOND) AS window_start,
service,
APPROX_PERCENTILE(latency_ms, 0.95) AS p95_latency
FROM input_stream
WHERE event_time >= CURRENT_TIMESTAMP - INTERVAL '2' HOUR
GROUP BY
service,
HOP(event_time, INTERVAL '15' SECOND, INTERVAL '1' MINUTE)
Step 4: Define alerting rules and smoothing
Set thresholds (e.g.: p95_latency > 500ms) and additional rules to avoid alerts from short spikes: require N consecutive windows above the threshold or use backoff. Example: alert only if 3 consecutive windows (with 15s slide) exceed 500ms.
// Simple logic to detect 3 consecutive windows in pseudo-SQL/streaming WITH p95_per_window AS (
-- result from step 3
)
SELECT service, window_start, p95_latency,
SUM(CASE WHEN p95_latency > 500 THEN 1 ELSE 0 END) OVER (PARTITION BY service ORDER BY window_start ROWS BETWEEN 2 PRECEDING AND CURRENT ROW) AS last3_count
FROM p95_per_window
WHERE last3_count = 3
Step 5: Emitting the alert and integrating with notifications
When the alert condition is satisfied, emit a message to an alerts topic or invoke a webhook/Azure Function that sends e-mail, SMS or creates a ticket. Include context: service, p95, timestamps and a sample of requests for diagnosis.
// Minimal example of payload for webhook {
"service": "checkout-api",
"alert": "p95_latency_exceeded",
"p95_latency": 652,
"window_start": "2026-09-30T12:00:00Z",
"samples": ["req-123","req-456"]
}
Verify the outcome
Confirm that the pipeline emits p95 per window and that alerts appear only when the smoothing rule is satisfied. Test with synthetic events: inject elevated latencies for 45s to try to exceed the 3-consecutive-windows criterion. Check engine logs, alert topics and webhook reception.
Conclusion
You have implemented a Real-Time Analytics flow that computes latency percentiles over sliding windows and triggers alerts with smoothing to avoid false positives. Next steps: tune thresholds per service, store history for analysis and integrate real-time dashboards. Tip: start with less sensitive metrics (p90) to reduce noise and only then move to p95/p99.