Ragfish Logo
Get In Touch
Ragfish Logo

Book a Demo

OpenAI Embeddings

OpenAIEmbedding provides the embedding model integration between OpenAI and the Ragfish Core Framework.

Embeddings are used to convert text into vector representations. These vectors allow Ragfish to perform semantic search over your knowledge base.

The integration includes:

The implementation is provided by the @ragfish/openai package and can be configured through Settings.embedModel.

Installation

Install the OpenAI integration package:

npm install @ragfish/openai

Make sure the Ragfish Core package is also installed:

npm install @ragfish/core

API Key

Store your OpenAI API key in an environment variable.

OPENAI_API_KEY=your-api-key

Do not hard-code API credentials in your application source code.

Basic Configuration

Import Settings from @ragfish/core and OpenAIEmbedding from @ragfish/openai.
import { Settings } from "@ragfish/core";
import { OpenAIEmbedding } from "@ragfish/openai";

Settings.embedModel = new OpenAIEmbedding({
 apiKey: process.env.OPENAI_API_KEY
});

Once configured, Ragfish can use the embedding model as part of the knowledge and retrieval pipeline.

Once configured, Ragfish components such as Chat can use the configured language model.

What Are Embeddings?

An embedding represents text as a numerical vector.

For example:

"How many employees work here?"
           |
           v
     Embedding Model
           |
           v
[0.021, -0.143, 0.782, ...]

The resulting vector represents the semantic meaning of the text.

Ragfish can compare these vectors to identify content that is semantically related to a user's question.

Embeddings in the RAG Pipeline

The embedding model is primarily involved in two stages.

Knowledge Ingestion

When knowledge is ingested, document chunks are converted into embeddings.

 Document
   |
 ▼
 Chunk
   |
 ▼
OpenAIEmbedding
   |
 ▼
Vector
   |
 ▼
Vector Store

User Query

When a user asks a question, the question can also be converted into an embedding.

User Question
     |
    ▼
OpenAIEmbedding
     |
    ▼
Query Vector
     |
    ▼
Vector Search

The retrieval system then uses the query vector to find relevant knowledge.

OpenAI Embeddings with Qdrant

A common Ragfish configuration combines OpenAIEmbedding with QdrantVectorStore.

OpenAIEmbedding
      |
     ▼
Generate Vectors
      |
     ▼
QdrantVectorStore
      |
     ▼
QdrantRetriever
      |
     ▼
Chat

Example:

import { Settings } from "@ragfish/core";
import { OpenAIEmbedding } from "@ragfish/openai";
import { QdrantVectorStore } from "@ragfish/qdrant";

Settings.embedModel = new OpenAIEmbedding({
 apiKey: process.env.OPENAI_API_KEY
});

const vectorStore = new QdrantVectorStore({
 // Qdrant configuration
});

Embedding During Ingestion

A typical knowledge ingestion flow looks like:

Knowledge Source
     |
    ▼
  Ingestion
     |
     ▼
  Chunking
     |
    ▼
OpenAIEmbedding
     |
    ▼
  Vectors
     |
    ▼
Vector Store

Each searchable chunk is represented as a vector before being stored.

Embedding During Retrieval

When a user asks a question, the query follows a similar process:

User Question
     |
    ▼
OpenAIEmbedding
     |
    ▼
Query Vector
     |
    ▼
Vector Store
     |
    ▼
Relevant Chunks

The retrieved chunks can then be passed to the language model to generate the final response.

Configuration

The basic configuration requires an OpenAI API key:

const embedding = new OpenAIEmbedding({
 apiKey: process.env.OPENAI_API_KEY
});

Settings.embedModel = embedding;

Additional embedding configuration should follow the options exposed by the version of @ragfish/openai installed in your project.

Refer to the package's TypeScript definitions and API Reference for the exact supported configuration options.

Complete Example

import { Settings, Chat } from "@ragfish/core";

import {
 OpenAILLM,
 OpenAIEmbedding
} from "@ragfish/openai";

import {
 QdrantVectorStore,
 QdrantRetriever
} from "@ragfish/qdrant";

Settings.llm = new OpenAILLM({
 apiKey: process.env.OPENAI_API_KEY
});

Settings.embedModel = new OpenAIEmbedding({
 apiKey: process.env.OPENAI_API_KEY
});

const vectorStore = new QdrantVectorStore({
 // Qdrant configuration
});

const retriever = new QdrantRetriever({
 vectorStore,
 collectionName: "knowledge"
});

const chat = new Chat({
 retriever
});

const response = await chat.message(
 "What is Ragfish?"
);

console.log(response);

Embedding Consistency

The same embedding configuration should be used consistently when indexing and querying a knowledge base.

Conceptually:

     Indexing
            |
▼ OpenAIEmbedding | ▼ Stored Vectors |
▼ OpenAIEmbedding |
▼ Query

Using incompatible embedding configurations can affect retrieval quality.

Best Practices

When using OpenAIEmbedding:

    1. Store the API key in environment variables.

    2. Configure the embedding model during application startup.

    3. Use the same embedding configuration consistently for a knowledge base.

    4. Keep embedding configuration separate from application business logic.

    5. Choose an embedding model appropriate for your application's content and retrieval requirements.

    6. Refer to the installed package's TypeScript definitions for supported options.

Next Steps

You now understand how OpenAI embeddings fit into the Ragfish retrieval pipeline.

Continue to Examples to see complete Ragfish applications that combine OpenAI, embeddings, vector stores, retrieval, and chat.