When inserting records into Vearch, the mandatory vector field must be supplied as a dense numerical representation of the original content. These vectors are not part of Vearch itself—they are generated externally using purpose-built models or algorithms. Below we explore how to obtain such embeddings from images, text, and audio, with code examples that can be adapted to production pipelines.
Image Embeddings with a Vision Transformer (ViT)
While convolutional networks are popular, a Vision Transformer can produce rich global representations. The following snippet uses a pretrained ViT from the timm library.
import timm
import torch
from PIL import Image
from torchvision import transforms
model = timm.create_model('vit_base_patch16_224', pretrained=True, num_classes=0)
model.eval()
preprocess = transforms.Compose([
transforms.Resize(256),
transforms.CenterCrop(224),
transforms.ToTensor(),
transforms.Normalize(mean=[0.5, 0.5, 0.5], std=[0.5, 0.5, 0.5]),
])
def image_to_vector(image_path):
img = Image.open(image_path).convert('RGB')
tensor = preprocess(img).unsqueeze(0)
with torch.no_grad():
features = model(tensor)
return features.squeeze().numpy()
# usage
vec = image_to_vector('photo.jpeg')
Text Embeddings via Sentence Transformers
A practical alternative to BERT token pooling is the SentenceTransformer library, which directly outputs fixed‑size sentence embeddings. This approach simplifies the pipeline.
from sentence_transformers import SentenceTransformer
encoder = SentenceTransformer('all-MiniLM-L6-v2')
def text_to_vector(raw_text):
embedding = encoder.encode(raw_text, convert_to_numpy=True)
return embedding
# usage
vec = text_to_vector("A query about vector search.")
Audio Feature Extraction using OpenL3
OpenL3 provides deep audio embeddings that combine self‑supervised learning with mel spectrograms, offering a modern alternative to handcrafted MFCCs.
import soundfile as sf
import openl3
def audio_to_vector(audio_path):
audio_data, sr = sf.read(audio_path, dtype='float32')
emb, _ = openl3.get_audio_embedding(audio_data, sr, content_type="music")
# Average over frames to obtain a single vector
return emb.mean(axis=0)
# usage
vec = audio_to_vector('sample.flac')
Pushing Vectors in to Vearch
Once you have the vector representation, submitting it is a straightforward REST call. Assume a space named multimedia inside the database gallery with an indexed field embedding.
curl -X POST "http://localhost:9001/gallery/multimedia" \
-H "Content-Type: application/json" \
-d '{"embedding": [0.04, -0.12, 0.78, ...]}'
The length and values of the vector must correspond to the dimension defined in the space schema. Generating consistent embeddings from the same model ensures meaningful nearest‑neighbour queries later.