Generating Embedding Vectors for Data Ingestion in Vearch

When inserting records into Vearch, the mandatory vector field must be supplied as a dense numerical representation of the original content. These vectors are not part of Vearch itself—they are generated externally using purpose-built models or algorithms. Below we explore how to obtain such embeddings from images, text, and audio, with code examples that can be adapted to production pipelines.

Image Embeddings with a Vision Transformer (ViT)

While convolutional networks are popular, a Vision Transformer can produce rich global representations. The following snippet uses a pretrained ViT from the timm library.

import timm
import torch
from PIL import Image
from torchvision import transforms

model = timm.create_model('vit_base_patch16_224', pretrained=True, num_classes=0)
model.eval()

preprocess = transforms.Compose([
    transforms.Resize(256),
    transforms.CenterCrop(224),
    transforms.ToTensor(),
    transforms.Normalize(mean=[0.5, 0.5, 0.5], std=[0.5, 0.5, 0.5]),
])

def image_to_vector(image_path):
    img = Image.open(image_path).convert('RGB')
    tensor = preprocess(img).unsqueeze(0)
    with torch.no_grad():
        features = model(tensor)
    return features.squeeze().numpy()

# usage
vec = image_to_vector('photo.jpeg')

Text Embeddings via Sentence Transformers

A practical alternative to BERT token pooling is the SentenceTransformer library, which directly outputs fixed‑size sentence embeddings. This approach simplifies the pipeline.

from sentence_transformers import SentenceTransformer

encoder = SentenceTransformer('all-MiniLM-L6-v2')

def text_to_vector(raw_text):
    embedding = encoder.encode(raw_text, convert_to_numpy=True)
    return embedding

# usage
vec = text_to_vector("A query about vector search.")

Audio Feature Extraction using OpenL3

OpenL3 provides deep audio embeddings that combine self‑supervised learning with mel spectrograms, offering a modern alternative to handcrafted MFCCs.

import soundfile as sf
import openl3

def audio_to_vector(audio_path):
    audio_data, sr = sf.read(audio_path, dtype='float32')
    emb, _ = openl3.get_audio_embedding(audio_data, sr, content_type="music")
    # Average over frames to obtain a single vector
    return emb.mean(axis=0)

# usage
vec = audio_to_vector('sample.flac')

Pushing Vectors in to Vearch

Once you have the vector representation, submitting it is a straightforward REST call. Assume a space named multimedia inside the database gallery with an indexed field embedding.

curl -X POST "http://localhost:9001/gallery/multimedia" \
  -H "Content-Type: application/json" \
  -d '{"embedding": [0.04, -0.12, 0.78, ...]}'

The length and values of the vector must correspond to the dimension defined in the space schema. Generating consistent embeddings from the same model ensures meaningful nearest‑neighbour queries later.

Tags: vearch vector database embedding extraction ViT sentence transformers

Posted on Sat, 10 Oct 2026 16:47:00 +0000 by toro04