Recursive character splitting
- Refer Here for the sample which explains how recursive chunking works
-
try by having one seperator and then the second
-
Refer Here for the adopted fix for the ncert book
-
Refer Here for clean up
Vector Storage
-
Challanges:
- Source Document Updates
-
Explore different index creation options
- Explore different retrieval mechanisms
Embedding storage layer: What gets stored
- Vectors (typically 384 – 3072 dim) are stored as dense float arrays
- Along side each vector, the followng are stored
- unique id
- metadata
- [optional] payload (raw text)
Indexing algorithm (core)
- lets assume we have around 50K chunks in vector database
- When you try to find a similar text
- Brute force cosine similarity is run on all the items in vector db
- Arranging indexes
- Flat
- IVF (Inverted File index)
- HNSW (Hierarchical Navigatable small world)
Distance Metric
- How closeness is actually computed
- Cosine Similarity
- Dot product
- Euclidean
Hybrid search
- Dense vs sparse Retrieval
Other Areas
- Metadata filtering
- CRUD/Updates
