AI-901: how to integrate Speech services in Microsoft Foundry
In this guide I will teach how to implement and integrate Speech services (speech recognition and voice synthesis) in Microsoft Foundry — a useful skill both for the AI-901 exam and for building conversational applications, automatic transcription and accessibility. We will focus on the essential concepts and a practical flow to connect Speech Services to Foundry pipelines.
What you need to know
Azure Speech Services include two main blocks: Speech-to-Text (STT) to transcribe speech to text, and Text-to-Speech (TTS) to synthesize voice from text. In the context of Microsoft Foundry, Foundry’s role is to orchestrate data and inference pipelines, including calls to managed Azure services or to local models. The integration enables, for example, ingestion of audio from multiple sources, processing (cleaning, normalization), transcription via Speech-to-Text, post-processing of the text (normalization, language detection) and, if desired, response synthesis with Text-to-Speech.
Practical example: imagine a pipeline that receives audio files recorded by call center agents. Foundry ingests those files, calls the Speech-to-Text service, stores the transcriptions in a OneLake table and automatically flags calls that contain certain keywords for further analysis.
How it works
The general steps to integrate Speech Services with Foundry are:
- Set up a Speech resource in Azure (key and endpoint) or use a compatible managed endpoint.
- In Foundry, create a pipeline that ingests the audio file (WAV/MP3) into OneLake or into a temporary dataset.
- Add an inference component that calls the Speech-to-Text endpoint, sending the audio or a URL to the file in OneLake.
- Receive the response (JSON with transcriptions, timestamps, confidence) and apply transformations: clean text, normalize punctuation, identify language.
- Save the transcription and metadata (duration, average confidence) in a table for reporting or trigger subsequent steps (e.g., sentiment analysis, call flagging).
- If you need to return audio (automated responses), call the Text-to-Speech service with the produced text and store/stream the output audio.
Example of a simplified payload for STT (sent by Foundry to an Azure Speech REST endpoint):
{
"audioUrl": "https://onelake/.../call123.wav",
"language": "pt-PT",
"properties": { "diarizationEnabled": true }
}
Typical response (simplified):
{
"transcription": "Olá, bom dia. Gostava de falar sobre a minha fatura.",
"confidence": 0.93,
"segments": [ { "start": 0.0, "end": 2.5, "text": "Olá, bom dia." }, ... ]
}
In practice (step-by-step inside Foundry)
1) Prepare credentials: add the Speech keys/endpoint as secrets in Foundry to avoid exposing credentials in code.
2) Ingestion: create a dataset that points to where the audio files arrive (Blob/OneLake). Normalize formats (use 16 kHz WAV mono when possible).
3) Inference component: use an HTTP operator or a native connector (if available) to call the Speech service. Define timeouts and retry policies; audio can be sent by URL or streamed.
4) Post-processing: apply transformations with scripts (Python/SQL) in the Foundry pipeline to clean the transcription — remove noise, normalize characters, apply case and automatic punctuation if needed.
5) Storage and action: write the transcription to a table and create triggers in the pipeline to execute additional analyses (e.g., intent detection, sentiment analysis, ticket generation).
6) Text-to-Speech (optional): when the flow requires an audio response, invoke the TTS endpoint with the final text and save the audio file for streaming to the user.
Common errors
1) Audio with unsupported format or sample rate: sending low-quality MP3s or unsupported sample rates reduces accuracy. Convert to 16 kHz WAV mono when possible.
2) Exposing credentials in the pipeline: placing keys directly in scripts or code in Foundry is a risk. Use Foundry’s secret manager and access controls.
3) Ignoring latency and cost management: frequent calls to Speech services can incur costs and latency. Batch files, use batch transcription for large files and adjust retry/timeout policies.
How to practice
To practice, use the OFFICIAL Microsoft Practice Assessment (free) to become familiar with the exam format. Also consult Microsoft’s study guide (free) on Azure AI Fundamentals and Speech Services to review concepts and official examples. Practically, build a small lab: create a Speech resource in your free Azure subscription, upload some audio files to OneLake or Blob Storage and build a simple pipeline in Foundry that calls Speech-to-Text and writes transcriptions. Avoid using sensitive real content during tests.
In summary
- Speech Services encompasses Speech-to-Text and Text-to-Speech; in Foundry it is used to transcribe audio and synthesize responses.
- Typical flow: audio ingestion → STT call → post-processing → storage/actions; TTS is optional for audio responses.
- Best practices: normalize audio format, use secrets for credentials and manage latency/costs with batching and retry policies.
- Practice with the official Practice Assessment and Microsoft’s study guide; build a lab with a Speech resource and a pipeline in Foundry.