How to generate multimodal embeddings in Azure AI & Machine Learning
Generating multimodal embeddings (text + image) allows comparing and searching heterogeneous content using a common representation. This tutorial shows how to create, test and use multimodal embeddings in Azure AI & Machine Learning for tasks such as semantic search or clustering.
Prerequisites
- Azure account with an active subscription and permissions to create resources.
- Azure OpenAI Service or Azure AI embeddings endpoint with multimodal support.
- Azure Machine Learning workspace (optional for experimentation and deployment).
- Python 3.9+ and pip; libraries: requests, numpy, scikit-learn, pillow.
- Sample set: some image files (.jpg/.png) and short texts for each item.
Step 1: Why multimodal embeddings?
Embeddings convert text and images into numeric vectors in the same space. This makes it possible to compare a text with an image (e.g. "red chair") or group similar items. In Azure AI & Machine Learning we use a multimodal embeddings endpoint to generate these vectors and then standard ML tools for indexing and search.
Step 2: Prepare the environment and credentials
Save the endpoint and key for the Azure OpenAI service or the multimodal embeddings service. Install the required libraries and prepare a folder with images and a CSV file with associated text.
python -m venv venv && source venv/bin/activate
pip install requests numpy scikit-learn pillow
export AZURE_EMBED_ENDPOINT="https://seu-endpoint.openai.azure.com"
export AZURE_EMBED_KEY="sua_chave"
Step 3: Minimal function to generate a text embedding
Call the embeddings endpoint for text. The example uses requests and assumes an endpoint compatible with the Azure OpenAI embeddings API.
import os, requests, numpy as np
ENDPOINT = os.environ['AZURE_EMBED_ENDPOINT']
KEY = os.environ['AZURE_EMBED_KEY']
HEADERS = {"api-key": KEY, "Content-Type": "application/json"}
def embed_text(text):
url = ENDPOINT + "/openai/deployments/embeddings/deploy-id/embeddings?api-version=2024-06-01"
data = {"input": text}
r = requests.post(url, headers=HEADERS, json=data)
r.raise_for_status()
vec = r.json()["data"][0]["embedding"]
return np.array(vec, dtype=float)
Step 4: Generate an image embedding (pre-processing)
Some multimodal endpoints accept an image as direct input; in other cases, we convert the image to base64 or extract local features with a lightweight model. Here we show how to send an image as base64 in case the endpoint accepts it.
import base64
from PIL import Image
import io
def embed_image(path):
with open(path, "rb") as f:
b64 = base64.b64encode(f.read()).decode('utf-8')
url = ENDPOINT + "/openai/deployments/embeddings-multimodal/deploy-id/embeddings?api-version=2024-06-01"
data = {"input": {"image": b64}}
r = requests.post(url, headers=HEADERS, json=data)
r.raise_for_status()
vec = r.json()["data"][0]["embedding"]
return np.array(vec, dtype=float)
Step 5: Build a simple index and perform semantic search
Store all vectors in an array and use cosine to measure similarity. Below is a simple example with scikit-learn to compute distances and return the top-k similar items.
from sklearn.metrics.pairwise import cosine_similarity
# Example mixed collection
items = [
{"id": "img1", "type": "image", "path": "images/cadeira_vermelha.jpg"},
{"id": "txt1", "type": "text", "text": "cadeira vermelha em madeira"},
]
vectors = []
meta = []
for it in items:
if it['type'] == 'image':
v = embed_image(it['path'])
else:
v = embed_text(it['text'])
vectors.append(v)
meta.append(it)
vectors = np.vstack(vectors)
def search(query_text, top_k=3):
qv = embed_text(query_text).reshape(1, -1)
sims = cosine_similarity(qv, vectors)[0]
idx = sims.argsort()[::-1][:top_k]
return [(meta[i], float(sims[i])) for i in idx]
# Search example
print(search("cadeira vermelha"))
Step 6: Common errors and how to resolve them
Common errors: 401/403 (wrong key or endpoint), 415 (unsupported image format), vectors with different dimensions (use the same deployment/model). Check the api-version and the deploy-id; if the endpoint does not support base64 images, use a local model to extract features before comparing.
Verify the result
Test with various text queries and compare the returned images/texts. The result is correct if semantically related items appear at the top. Also check that the embeddings have a fixed dimension and that cosine similarity returns values close to 1 for identical items.
Conclusion
With multimodal embeddings in Azure AI & Machine Learning you can unite text and image in the same vector space for semantic search, clustering or classification. Next steps: index vectors with FAISS or Azure Cognitive Search for production, and experiment with similarity thresholds. Tip: start with a small dataset and verify the embedding dimensions before scaling.