Chunking
Chunking is the process of dividing a document or other source content into smaller units that can be processed, embedded, stored, and retrieved efficiently.
In a RAG application, the quality of chunking has a direct impact on retrieval quality. If chunks are too large, they may contain too much unrelated information. If they are too small, important context may be lost.
Ragfish treats chunking as a separate stage in the knowledge pipeline so that it can be configured independently from ingestion, embeddings, and retrieval.
Where Chunking Fits
Chunking sits between ingestion and embedding in the Ragfish knowledge pipeline.
Knowledge Source
│
▼
Ingestion
│
▼
Document
│
▼
Chunking
│
▼
Chunks
│
▼
Embeddings
│
▼
Vector Store
The ingestion layer prepares the source content, while the chunking layer determines how that content is divided.
Why Chunking Matters
Consider a large document containing several topics:
Company Handbook
Introduction
Employee Benefits
Leave Policy
Travel Policy
Work From Home Policy
Security Policy
If the entire document is treated as one searchable unit, a question about the leave policy may retrieve a large amount of unrelated content.
With appropriate chunking:
Document │ ├── Introduction ├── Employee Benefits ├── Leave Policy ├── Travel Policy ├── Work From Home Policy └── Security Policy
The retrieval system can identify the relevant section more precisely.
What Is a Chunk?
A chunk is a smaller unit of content created from a larger document.
A chunk can contain:
Text
Metadata
Source information
Document context
Conceptually:
Document
│
├── Chunk 1
│ ├── Content
│ └── Metadata
│
├── Chunk 2
│ ├── Content
│ └── Metadata
│
└── Chunk 3
├── Content
└── Metadata
The exact structure of a chunk depends on the Ragfish implementation and the source being processed.
Chunking and Embeddings
Once content has been divided into chunks, each chunk can be converted into an embedding.
Document
│
▼
Chunking
│
├── Chunk 1 ──► Embedding 1
├── Chunk 2 ──► Embedding 2
└── Chunk 3 ──► Embedding 3
│
▼
Vector Store
The resulting vectors represent the semantic meaning of the individual chunks.
During retrieval, the user's question is also represented as a vector and compared against the stored embeddings.
Chunk Size
Chunk size determines how much content is included in each chunk.
For example:
Small Chunks ──────────────── Short sections Precise retrieval Less surrounding context Large Chunks ──────────────── More context More information per result Potentially more unrelated content
There is no single chunk size that works for every application.
The appropriate size depends on:
Document structure
Content type
Question patterns
Embedding model
Retrieval strategy
Expected response context
Chunk Overlap
When documents are split into chunks, information at the boundary between two chunks can sometimes be lost.
Chunk overlap helps preserve that context.
For example:
Without Overlap Chunk 1 [Introduction ........ Policy begins] Chunk 2 [Policy details ........ Conclusion] With overlap: With Overlap Chunk 1 [Introduction .... Policy begins ....] Chunk 2 [.... Policy begins .... Policy details ....] Chunk 3 [.... Policy details .... Conclusion]
The overlapping content provides continuity between neighboring chunks.
Structure-Aware Chunking
Not all documents should be treated as plain text.
Structured documents often contain meaningful boundaries such as:
Headings
Sections
Paragraphs
Tables
Lists
Pages
Whenever possible, chunking should preserve meaningful document structure.
For example:
Employee Handbook
│
├── Leave
│ ├── Annual Leave
│ ├── Sick Leave
│ └── Holiday Policy
│
└── Benefits
├── Health Insurance
└── Retirement
Preserving these relationships can improve the context available during retrieval.
Chunk Metadata
Metadata provides additional information about where a chunk came from.
For example:
Chunk ├── Content ├── Document ├── Page ├── Section └── Source
Metadata can later help the retrieval layer identify, filter, or display the source of retrieved information.
This is particularly useful for enterprise applications where users may need to understand where an answer originated.
Chunking Different Data Sources
Different knowledge sources may require different approaches.
Documents
Documents may benefit from section- or paragraph-aware chunking.
Spreadsheets
Spreadsheet data may need to preserve relationships between columns, rows, headers, and records.
Websites
Web content may need to preserve page structure, headings, and sections.
Databases
Database content may need to retain relationships between records and fields.
The chunking strategy should therefore reflect the structure of the underlying source.
Chunking and Retrieval
Chunking and retrieval are closely connected.
The chunking process determines the units that retrieval can search.
Document
│
▼
Chunking
│
┌───────────┼───────────┐
▼ ▼ ▼
Chunk 1 Chunk 2 Chunk 3
│ │ │
└───────────┼───────────┘
▼
Embeddings
│
▼
Vector Store
│
▼
Retrieval
Poor chunking can make retrieval less effective even when the vector database and embedding model are configured correctly.
Choosing a Chunking Strategy
When designing a chunking strategy, consider:
Document Structure
Does the source have meaningful headings, sections, or records?
Content Length
Long documents may require more aggressive chunking.
Question Type
If users ask highly specific questions, smaller and more focused chunks may work better.
Context Requirements
If understanding an answer requires surrounding context, chunks may need to be larger or overlap with neighboring content.
Metadata
Preserve metadata that can help identify and filter retrieved content.
Custom Chunking
Different applications may require different chunking strategies.
Ragfish's modular architecture is designed so that chunking can remain independent from the rest of the RAG pipeline.
Conceptually:
Ingestion
│
▼
Chunking Strategy
│
▼
Embeddings
│
▼
Vector Store
│
▼
Retrieval
This allows an application to change how documents are divided without redesigning the entire retrieval system.
Detailed custom chunking APIs should follow the chunking interfaces implemented in the installed Ragfish version.
Common Chunking Problems
Chunks Are Too Large
Large chunks may contain unrelated information and reduce retrieval precision.
Chunks Are Too Small
Very small chunks may lose the context required to understand the information.
Important Boundaries Are Lost
Splitting content in the middle of a section, table, or logical unit can reduce retrieval quality.
Metadata Is Lost
Removing source information makes it harder to identify where retrieved information originated.
Best Practices
When designing chunking strategies:
Preserve meaningful document boundaries.
Choose chunk sizes based on the content.
Use overlap when context spans chunk boundaries.
Preserve useful metadata.
Treat structured sources differently from plain text.
Test chunking with real user questions.
Evaluate retrieval quality after changing chunking settings.
The best chunking strategy is the one that consistently produces relevant context for your application's real queries.
Chunking in the Complete RAG Pipeline
Chunking is one stage of the complete Ragfish knowledge pipeline:
Knowledge Source
│
▼
Ingestion
│
▼
Chunking
│
▼
Embeddings
│
▼
Vector Store
│
▼
Retrieval
│
▼
Chat
│
▼
AI Assistant
Each stage has a specific responsibility. Keeping these stages separate makes the Ragfish Framework easier to customize and extend.
What's Next?
Chunks are created so that the retrieval system can find the most relevant knowledge for a user's question.
Continue to Retrieval to learn how Ragfish searches stored knowledge, selects relevant chunks, and provides context to the language model.