(+351) 21 24 10006  ·  info@bconcepts.pt
Carnaxide, Lisbon

How to generate multimodal embeddings in Azure AI & Machine Learning

João Barros 03 de October de 2026 4 min read

Generating multimodal embeddings (text + image) allows comparing and searching heterogeneous content using a common representation. This tutorial shows how to create, test and use multimodal embeddings in Azure AI & Machine Learning for tasks such as semantic search or clustering.

Prerequisites

  • Azure account with an active subscription and permissions to create resources.
  • Azure OpenAI Service or Azure AI embeddings endpoint with multimodal support.
  • Azure Machine Learning workspace (optional for experimentation and deployment).
  • Python 3.9+ and pip; libraries: requests, numpy, scikit-learn, pillow.
  • Sample set: some image files (.jpg/.png) and short texts for each item.

Step 1: Why multimodal embeddings?

Embeddings convert text and images into numeric vectors in the same space. This makes it possible to compare a text with an image (e.g. "red chair") or group similar items. In Azure AI & Machine Learning we use a multimodal embeddings endpoint to generate these vectors and then standard ML tools for indexing and search.

Step 2: Prepare the environment and credentials

Save the endpoint and key for the Azure OpenAI service or the multimodal embeddings service. Install the required libraries and prepare a folder with images and a CSV file with associated text.

python -m venv venv && source venv/bin/activate
pip install requests numpy scikit-learn pillow
export AZURE_EMBED_ENDPOINT="https://seu-endpoint.openai.azure.com"
export AZURE_EMBED_KEY="sua_chave"

Step 3: Minimal function to generate a text embedding

Call the embeddings endpoint for text. The example uses requests and assumes an endpoint compatible with the Azure OpenAI embeddings API.

import os, requests, numpy as np

ENDPOINT = os.environ['AZURE_EMBED_ENDPOINT']
KEY = os.environ['AZURE_EMBED_KEY']
HEADERS = {"api-key": KEY, "Content-Type": "application/json"}

def embed_text(text):
    url = ENDPOINT + "/openai/deployments/embeddings/deploy-id/embeddings?api-version=2024-06-01"
    data = {"input": text}
    r = requests.post(url, headers=HEADERS, json=data)
    r.raise_for_status()
    vec = r.json()["data"][0]["embedding"]
    return np.array(vec, dtype=float)

Step 4: Generate an image embedding (pre-processing)

Some multimodal endpoints accept an image as direct input; in other cases, we convert the image to base64 or extract local features with a lightweight model. Here we show how to send an image as base64 in case the endpoint accepts it.

import base64
from PIL import Image
import io

def embed_image(path):
    with open(path, "rb") as f:
        b64 = base64.b64encode(f.read()).decode('utf-8')
    url = ENDPOINT + "/openai/deployments/embeddings-multimodal/deploy-id/embeddings?api-version=2024-06-01"
    data = {"input": {"image": b64}}
    r = requests.post(url, headers=HEADERS, json=data)
    r.raise_for_status()
    vec = r.json()["data"][0]["embedding"]
    return np.array(vec, dtype=float)

Step 5: Build a simple index and perform semantic search

Store all vectors in an array and use cosine to measure similarity. Below is a simple example with scikit-learn to compute distances and return the top-k similar items.

from sklearn.metrics.pairwise import cosine_similarity

# Example mixed collection
items = [
    {"id": "img1", "type": "image", "path": "images/cadeira_vermelha.jpg"},
    {"id": "txt1", "type": "text", "text": "cadeira vermelha em madeira"},
]

vectors = []
meta = []
for it in items:
    if it['type'] == 'image':
        v = embed_image(it['path'])
    else:
        v = embed_text(it['text'])
    vectors.append(v)
    meta.append(it)

vectors = np.vstack(vectors)

def search(query_text, top_k=3):
    qv = embed_text(query_text).reshape(1, -1)
    sims = cosine_similarity(qv, vectors)[0]
    idx = sims.argsort()[::-1][:top_k]
    return [(meta[i], float(sims[i])) for i in idx]

# Search example
print(search("cadeira vermelha"))

Step 6: Common errors and how to resolve them

Common errors: 401/403 (wrong key or endpoint), 415 (unsupported image format), vectors with different dimensions (use the same deployment/model). Check the api-version and the deploy-id; if the endpoint does not support base64 images, use a local model to extract features before comparing.

Verify the result

Test with various text queries and compare the returned images/texts. The result is correct if semantically related items appear at the top. Also check that the embeddings have a fixed dimension and that cosine similarity returns values close to 1 for identical items.

Conclusion

With multimodal embeddings in Azure AI & Machine Learning you can unite text and image in the same vector space for semantic search, clustering or classification. Next steps: index vectors with FAISS or Azure Cognitive Search for production, and experiment with similarity thresholds. Tip: start with a small dataset and verify the embedding dimensions before scaling.