Ingestion
Ingestion is the process of bringing external knowledge into the Ragfish Framework so that it can be processed, indexed, and retrieved by an AI application.
Ragfish is designed to work with knowledge from different sources, including documents, spreadsheets, databases, and other connectors.
The ingestion layer provides the bridge between these external knowledge sources and the retrieval system.
The Ingestion Pipeline
A typical Ragfish ingestion pipeline can be represented as:
Knowledge Source
│
▼
Connector
│
▼
Ingestion
│
▼
Document Data
│
▼
Chunking
│
▼
Embeddings
│
▼
Vector Store
The purpose of ingestion is to transform source data into a format that can be used efficiently by the retrieval system.
What Does Ingestion Do?
Depending on the connector and source type, the ingestion process can involve several stages.
1. Load
Read data from the configured knowledge source.
Examples include:
Documents
Spreadsheets
Databases
Websites
Documentation
2. Extract
Extract the relevant content from the source.
For example, a document connector may extract text while a spreadsheet connector may read rows, columns, and cell values.
3. Prepare
Prepare the extracted content for downstream processing.
This can include:
Normalizing content
Preserving metadata
Identifying document boundaries
Preparing content for chunking
4. Chunk
Large content is divided into smaller units that can be processed and retrieved efficiently.
Chunking is covered in detail in the Chunking section.
5. Embed
The processed chunks can be converted into vector representations using the configured embedding model.
6. Store
The resulting vectors and associated information are stored in the configured vector store.
Connectors and Ingestion
Ragfish separates where knowledge comes from from how knowledge is processed.
A connector is responsible for accessing a particular source.
For example:
Spreadsheet
│
▼
Spreadsheet Connector
│
▼
Ingestion Pipeline
│
▼
Chunks
│
▼
Embeddings
│
▼
Vector Store
This separation allows the framework to support different knowledge sources without changing the core ingestion architecture.
Supported Knowledge Sources
The Ragfish ecosystem is designed around a connector-based architecture.
Current and planned sources include:
Spreadsheet
PDF
Database
Website
Documentation
Notion
SharePoint
Google Drive
The availability of each connector depends on the current Ragfish release.
Ingestion and Metadata
Knowledge is often more useful when the original context is preserved.
For example, a document may contain metadata such as:
Document ├── Title ├── Source ├── File Name ├── Page ├── Section └── Content
Metadata can later be used during retrieval to identify where a piece of information originated or to apply filtering rules.
The exact metadata model depends on the connector and framework implementation.
Ingestion and Chunking
Ingestion and chunking are closely related, but they have different responsibilities.
Ingestion brings knowledge into the framework.
Chunking determines how that knowledge is divided into smaller searchable units.
Source │ ▼ Ingestion │ ▼ Document │ ▼ Chunking │ ├── Chunk 1 ├── Chunk 2 ├── Chunk 3 └── Chunk 4
Keeping these responsibilities separate makes it possible to change chunking strategies without changing the source connector.
Ingestion and Embeddings
After content has been prepared and chunked, the chunks can be converted into embeddings.
Document
│
▼
Chunks
│
▼
Embedding Model
│
▼
Vectors
│
▼
Vector Store
The embedding model is configured through the Ragfish framework settings.
For example:
Settings.embedModel = new OpenAIEmbedding({
apiKey: process.env.OPENAI_API_KEY
});
The embedding provider is independent from the ingestion source.
This means the same ingestion architecture can work with different embedding providers.
Ingestion and Retrieval
Ingestion prepares knowledge before a user asks a question.
Retrieval uses that prepared knowledge when a user asks a question.
INGESTION
│
▼
Prepare Knowledge
│
▼
Vector Store
│
│
▼
RETRIEVAL
│
User Question
│
▼
Relevant Knowledge
│
▼
LLM
This separation is fundamental to a RAG application.
Example:Spreadsheet Ingestion
Consider an Excel file containing product information.
products.xlsx
The ingestion flow can be represented as:
Excel File
│
▼
Spreadsheet Connector
│
▼
Extract Spreadsheet Data
│
▼
Prepare Content
│
▼
Chunk Data
│
▼
Generate Embeddings
│
▼
Store in Vector Database
Once the data has been indexed, an assistant can retrieve relevant information when a user asks a question.
For example:
"How many methods can ADM3 store?"
The retriever searches the indexed knowledge and provides the relevant context to the language model. The Ragfish master architecture uses this same pattern with a spreadsheet/knowledge source, embedding model, vector store, retriever, and Chat.
Re-ingestion
Knowledge sources can change over time.
When source content changes, the application may need to process the updated information again.
A typical lifecycle is:
Source Updated
│
▼
Detect Changes
│
▼
Re-ingest Content
│
▼
Re-process Chunks
│
▼
Update Embeddings
│
▼
Update Vector Store
The exact update and synchronization strategy depends on the connector and application implementation.
Designing an Ingestion Pipeline
When designing an ingestion pipeline, consider:
Source type
Document size
Content structure
Metadata requirements
Chunking strategy
Embedding model
Vector store
Update frequency
Duplicate handling
Error handling
These decisions directly affect retrieval quality and application performance.
Best Practices
For reliable ingestion:
Keep source-specific logic inside connectors.
Preserve useful metadata during ingestion.
Use a chunking strategy appropriate for the content.
Keep embedding configuration independent from source configuration.
Plan for changes to source data.
Handle failed documents without stopping the entire ingestion process.
Keep the ingestion pipeline reproducible.
Ingestion Architecture
The complete relationship between connectors, ingestion, chunking, embeddings, and storage can be summarized as:
Knowledge Sources
│
┌────────────┼────────────┐
▼ ▼ ▼
Spreadsheet PDF Database
│ │ │
└────────────┼────────────┘
▼
Connectors
│
▼
Ingestion
│
▼
Chunking
│
▼
Embeddings
│
▼
Vector Store
│
▼
Retrieval
│
▼
Chat
│
▼
AI Assistant
This architecture allows Ragfish to support multiple knowledge sources while maintaining a consistent retrieval experience.
What's Next?
Ingestion prepares knowledge for the retrieval system. The next step is understanding how that knowledge is divided into smaller, searchable units.
Continue to Chunking to learn how Ragfish processes documents into chunks, how chunk size and structure affect retrieval, and how to choose an appropriate chunking strategy for your application.