Implementing Semantic Search and Data Visualization with OpenAI Embeddings

Foundations of Knowledge Representation

The advancement of artificial intelligence relies heavily on how machines represent and process objective knowledge. While defining intelligence remains abstract, the Turing Test suggests we evaluate machine intelligence by its ability to exhibit indistinguishable behavior from humans. In digital systems, knowledge is encoded using specific standards. For Western languages, ASCII (American Standard Code for Information Interchange) serves as a foundational mapping between characters and numerical values. For languages with extensive character sets like Chinese, more complex encoding systems are utilized to represent the vast array of logograms efficiently.

Representation Learning and Vector Embeddings

Representation learning involves algorithms that automatically discover the most effective ways to represent raw data for machine learning tasks. The goal is to transform input data into a feature space that highlights essential structures, making the data more separable, interpretable, and suitable for reasoning.

Embeddings are a specific implementation of representation learning where high-dimensional data is mapped into a lower-dimensional vector space. In Natural Language Processing (NLP), word embeddings convert discrete words into continuous vectors. This mapping ensures that words with semantic similarities are positioned closer together in the vector space, while unrelated words are further apart.

Key Objectives of Representation Learning

  • Separability: Transforming data so that different categories can be easily distinguished by a classifier.
  • Interpretability: Creating features that correspond to understandable human concepts, aiding in model debugging and trust.
  • Reasoning Capabilities: Enabling arithmetic operations on vectors to infer relationships. For example, the vector operation King - Man + Woman results in a vector close to Queen.

Visualizing High-Dimensional Data

Since embeddings often exist in high-dimensional spaces (e.g., 1536 dimensions), visualization requires dimensionality reduction. t-SNE (t-Distributed Stochastic Neighbor Embedding) is a statistical technique well-suited for this. It projects high-dimensional data into 2D or 3D space, preserving the local neighborhood structure. This allows developers to visually inspect clusters and identify patterns that are not apparent in raw data.

Applications of Word Embeddings

Word embeddings provide robust semantic representations for text:

  • Semantic Similarity: Measuring the proximity of word vectors to determine if words are related.
  • Relationship Extraction: Identifying analogies and syntactic relationships (e.g., verb tenses) through vector offsets.
  • Contextual Disambiguation: Helping to resolve word meaning based on surrounding context by using vector representations.
  • Downstream Task Enhancement: Improving performance in text classification, sentiment analysis, and machine translation by providing dense, meaningful input features.

OpenAI Embeddings: Practical Implementation

Environment Configuration

First, install the necessary Python libraries and configure the API key.

pip install openai tiktoken pandas numpy scikit-learn matplotlib

Ensure the OPENAI_API_KEY environment variable is set correctly for authentication.

Dataset Preparation

We will utilize the Amazon Fine Food Reviews dataset. For demonstration purposes, we will process a subset of 1,000 recent reviews. The goal is to convert the review text into vector embeddings to enable semantic analysis.

import pandas as pd
import tiktoken
from openai.embeddings_utils import get_embedding

# Load the dataset
data_path = "fine_food_reviews_1k.csv"
df = pd.read_csv(data_path, index_col=0)
df = df[['Score', 'Summary', 'Text']].dropna()

# Combine the review summary and body
df['combined_text'] = df['Summary'].str.strip() + "; " + df['Text'].str.strip()

# Define model parameters
model_name = "text-embedding-ada-002"
encoding_name = "cl100k_base"
token_limit = 8000

# Filter out reviews exceeding the token limit
encoding = tiktoken.get_encoding(encoding_name)
df['token_count'] = df['combined_text'].apply(lambda x: len(encoding.encode(x)))
df = df[df['token_count'] <= token_limit].tail(1000)

Generating Embeddings

We generate the embeddings using OpenAI's text-embedding-ada-002 model, which outputs a vector with 1536 dimensions for each input text.

# Generate vector embeddings
df['vector_embedding'] = df['combined_text'].apply(lambda x: get_embedding(x, engine=model_name))

# Save the processed data
df.to_csv("food_reviews_with_embeddings.csv")

Visualization with t-SNE and K-Means

To visualize the distribution of these reviews, we reduce the dimensionality to 2D using t-SNE and apply K-Means clustering to group similar reviews.

import numpy as np
import matplotlib.pyplot as plt
from sklearn.manifold import TSNE
from sklearn.cluster import KMeans
import matplotlib

# Convert embeddings to a matrix
matrix = np.vstack(df['vector_embedding'].values)

# Apply t-SNE for dimensionality reduction
tsne = TSNE(n_components=2, perplexity=15, random_state=42, init='random')
vis_dims = tsne.fit_transform(matrix)

# Visualize by Review Score
x_vals = [x for x, y in vis_dims]
y_vals = [y for x, y in vis_dims]
score_colors = ["red", "orange", "gold", "turquoise", "green"]
color_indices = df['Score'].values - 1

plt.figure(figsize=(10, 7))
plt.scatter(x_vals, y_vals, c=color_indices, cmap=matplotlib.colors.ListedColormap(score_colors), alpha=0.3)
plt.title("Amazon Food Reviews Visualized by Score")
plt.show()

# Apply K-Means Clustering
num_clusters = 4
kmeans = KMeans(n_clusters=num_clusters, init='k-means++', random_state=42, n_init=10)
df['cluster_group'] = kmeans.fit_predict(matrix)

# Visualize Clusters
cluster_colors = ["purple", "blue", "green", "yellow"]
plt.figure(figsize=(10, 7))
plt.scatter(x_vals, y_vals, c=df['cluster_group'], cmap=matplotlib.colors.ListedColormap(cluster_colors), alpha=0.3)
plt.title("Review Clusters Visualized using t-SNE")
plt.show()

Implementing Semantic Search

Embeddings allow for semantic search by calculating the cosine similarity between a query vector and the document vectors.

from openai.embeddings_utils import cosine_similarity

def find_similar_reviews(dataframe, search_query, n=3):
    # Convert the query to an embedding
    query_embedding = get_embedding(search_query, engine=model_name)
    
    # Calculate cosine similarity for all reviews
    dataframe['similarity'] = dataframe['vector_embedding'].apply(
        lambda x: cosine_similarity(x, query_embedding)
    )
    
    # Return the top n most similar reviews
    results = dataframe.sort_values('similarity', ascending=False).head(n)
    return results['combined_text']

# Execute a search
results = find_similar_reviews(df, "delicious coffee beans", n=3)
for idx, text in enumerate(results):
    print(f"Match {idx+1}: {text[:150]}...")

Assessment Questions

  1. What is the primary goal of representation learning?
    A. To increase data size
    B. To transform raw data into a more useful feature representation
    C. To generate random data
    D. To simplify data into a single value
  2. In Natural Language Processing, which data format is most commonly processed?
    A. Images
    B. Text
    C. Audio
    D. Video
  3. What type of model is Word2Vec?
    A. An image generator
    B. A word representation learning model
    C. A speech synthesizer
    D. A database manager
  4. Which method is commonly used to measure the similarity between two word vectors?
    A. Euclidean distance
    B. Cosine similarity
    C. Hamming distance
    D. Levenshtein distance
  5. Why is unsupervised learning frequently applied in representation learning?
    A. It is always faster
    B. It leverages unlabeled data which is abundant
    C. It requires no computational resources
    D. It guarantees perfect accuracy

Tags: OpenAI Embeddings NLP Machine Learning python

Posted on Fri, 02 Oct 2026 16:48:40 +0000 by codygoodman