Ragfish Logo
Get In Touch
Ragfish Logo

Book a Demo

Ingestion

Ingestion is the process of bringing external knowledge into the Ragfish Framework so that it can be processed, indexed, and retrieved by an AI application.

Ragfish is designed to work with knowledge from different sources, including documents, spreadsheets, databases, and other connectors.

The ingestion layer provides the bridge between these external knowledge sources and the retrieval system.

The Ingestion Pipeline

A typical Ragfish ingestion pipeline can be represented as:

Knowledge Source
      │
      ▼
   Connector
      │
      ▼
   Ingestion
      │
      ▼
  Document Data
      │
      ▼
   Chunking
      │
      ▼
  Embeddings
      │
      ▼
 Vector Store

The purpose of ingestion is to transform source data into a format that can be used efficiently by the retrieval system.

What Does Ingestion Do?

Depending on the connector and source type, the ingestion process can involve several stages.

1. Load

Read data from the configured knowledge source.

Examples include:

    1. Documents

    2. Spreadsheets

    3. Databases

    4. Websites

    5. Documentation

2. Extract

Extract the relevant content from the source.

For example, a document connector may extract text while a spreadsheet connector may read rows, columns, and cell values.

3. Prepare

Prepare the extracted content for downstream processing.

This can include:

    1. Normalizing content

    2. Preserving metadata

    3. Identifying document boundaries

    4. Preparing content for chunking

4. Chunk

Large content is divided into smaller units that can be processed and retrieved efficiently.

Chunking is covered in detail in the Chunking section.

5. Embed

The processed chunks can be converted into vector representations using the configured embedding model.

6. Store

The resulting vectors and associated information are stored in the configured vector store.

Connectors and Ingestion

Ragfish separates where knowledge comes from from how knowledge is processed.

A connector is responsible for accessing a particular source.

For example:

Spreadsheet
     │
     ▼
Spreadsheet Connector
     │
     ▼
Ingestion Pipeline
     │
     ▼
Chunks
     │
     ▼
Embeddings
     │
     ▼
Vector Store

This separation allows the framework to support different knowledge sources without changing the core ingestion architecture.

Supported Knowledge Sources

The Ragfish ecosystem is designed around a connector-based architecture.

Current and planned sources include:

    1. Spreadsheet

    2. PDF

    3. Database

    4. Website

    5. Documentation

    6. Notion

    7. SharePoint

    8. Google Drive

The availability of each connector depends on the current Ragfish release.

Ingestion and Metadata

Knowledge is often more useful when the original context is preserved.

For example, a document may contain metadata such as:

Document
├── Title
├── Source
├── File Name
├── Page
├── Section
└── Content

Metadata can later be used during retrieval to identify where a piece of information originated or to apply filtering rules.

The exact metadata model depends on the connector and framework implementation.

Ingestion and Chunking

Ingestion and chunking are closely related, but they have different responsibilities.

Ingestion brings knowledge into the framework.

Chunking determines how that knowledge is divided into smaller searchable units.

Source
  │
  ▼
Ingestion
  │
  ▼
Document
  │
  ▼
Chunking
  │
  ├── Chunk 1
  ├── Chunk 2
  ├── Chunk 3
  └── Chunk 4

Keeping these responsibilities separate makes it possible to change chunking strategies without changing the source connector.

Ingestion and Embeddings

After content has been prepared and chunked, the chunks can be converted into embeddings.

Document
    │
    ▼
  Chunks
    │
    ▼
Embedding Model
    │
    ▼
Vectors
    │
    ▼
Vector Store

The embedding model is configured through the Ragfish framework settings.

For example:

Settings.embedModel = new OpenAIEmbedding({
  apiKey: process.env.OPENAI_API_KEY
});

The embedding provider is independent from the ingestion source.

This means the same ingestion architecture can work with different embedding providers.

Ingestion and Retrieval

Ingestion prepares knowledge before a user asks a question.

Retrieval uses that prepared knowledge when a user asks a question.

 INGESTION
              │
              ▼
        Prepare Knowledge
              │
              ▼
         Vector Store
              │
              │
              ▼
          RETRIEVAL
              │
        User Question
              │
              ▼
       Relevant Knowledge
              │
              ▼
             LLM

This separation is fundamental to a RAG application.

Example:Spreadsheet Ingestion

Consider an Excel file containing product information.

products.xlsx

The ingestion flow can be represented as:

Excel File
    │
    ▼
Spreadsheet Connector
    │
    ▼
Extract Spreadsheet Data
    │
    ▼
Prepare Content
    │
    ▼
Chunk Data
    │
    ▼
Generate Embeddings
    │
    ▼
Store in Vector Database

Once the data has been indexed, an assistant can retrieve relevant information when a user asks a question.

For example:

"How many methods can ADM3 store?"

The retriever searches the indexed knowledge and provides the relevant context to the language model. The Ragfish master architecture uses this same pattern with a spreadsheet/knowledge source, embedding model, vector store, retriever, and Chat.

Re-ingestion

Knowledge sources can change over time.

When source content changes, the application may need to process the updated information again.

A typical lifecycle is:

Source Updated
      │
      ▼
Detect Changes
      │
      ▼
Re-ingest Content
      │
      ▼
Re-process Chunks
      │
      ▼
Update Embeddings
      │
      ▼
Update Vector Store

The exact update and synchronization strategy depends on the connector and application implementation.

Designing an Ingestion Pipeline

When designing an ingestion pipeline, consider:

    1. Source type

    2. Document size

    3. Content structure

    4. Metadata requirements

    5. Chunking strategy

    6. Embedding model

    7. Vector store

    8. Update frequency

    9. Duplicate handling

    10. Error handling

These decisions directly affect retrieval quality and application performance.

Best Practices

For reliable ingestion:

    1. Keep source-specific logic inside connectors.

    2. Preserve useful metadata during ingestion.

    3. Use a chunking strategy appropriate for the content.

    4. Keep embedding configuration independent from source configuration.

    5. Plan for changes to source data.

    6. Handle failed documents without stopping the entire ingestion process.

    7. Keep the ingestion pipeline reproducible.

Ingestion Architecture

The complete relationship between connectors, ingestion, chunking, embeddings, and storage can be summarized as:

               Knowledge Sources
                       │
          ┌────────────┼────────────┐
          ▼            ▼            ▼
     Spreadsheet      PDF       Database
          │            │            │
          └────────────┼────────────┘
                       ▼
                   Connectors
                       │
                       ▼
                   Ingestion
                       │
                       ▼
                    Chunking
                       │
                       ▼
                  Embeddings
                       │
                       ▼
                  Vector Store
                       │
                       ▼
                   Retrieval
                       │
                       ▼
                      Chat
                       │
                       ▼
                  AI Assistant

This architecture allows Ragfish to support multiple knowledge sources while maintaining a consistent retrieval experience.

What's Next?

Ingestion prepares knowledge for the retrieval system. The next step is understanding how that knowledge is divided into smaller, searchable units.

Continue to Chunking to learn how Ragfish processes documents into chunks, how chunk size and structure affect retrieval, and how to choose an appropriate chunking strategy for your application.