Default RAG Stack Series — Part 3: Embeddings & Vector Databases
How Meaning Becomes Math — and How We Store It Efficiently
Technology: Python, LangChain, Qdrant, SentenceTransformers
Skill: Vector Search, Semantic Similarity, Embedding Models, ANN Indexing
Default RAG Stack Series — Part 3
Embeddings & Vector Databases
How Meaning Becomes Math — and How We Store It Efficiently
1. From Words to Vectors: The Core Idea
Language models can’t “understand” words — they understand relationships between words.
Embeddings are the bridge between human language and machine reasoning.
They turn text into numerical vectors where semantic similarity translates to spatial proximity in high-dimensional space.
Formally, an embedding function maps text t to a vector:
$$ f_\theta:\text{text} \rightarrow \mathbb{R}^n $$
If two texts are similar in meaning, their vectors are close together:
$$ \text{sim}(t_1, t_2) = \frac{f_\theta(t_1) \cdot f_\theta(t_2)}{|f_\theta(t_1)||f_\theta(t_2)|} \approx 1 $$
This is the mathematical foundation that enables RAG retrieval — you can find relevant passages by geometry, not keywords.
2. Why Embeddings Matter in RAG
The retriever in RAG uses embeddings to find text fragments most semantically aligned with a user query.
Without them, retrieval would depend only on surface-level keyword overlap (like BM25).
| Retrieval Type | Strength | Limitation |
|---|---|---|
| Lexical (BM25, TF-IDF) | Exact word matching | Misses paraphrases |
| Semantic (Embeddings) | Captures meaning similarity | Requires vector infrastructure |
| Hybrid | Combines both | Slightly higher latency, better recall |
Embeddings make retrieval context-aware — “revenue grew 5%” and “sales increased by five percent” live close together in vector space even though they share no identical words.
3. Anatomy of an Embedding Vector
A typical sentence embedding is a vector of dimension n ∈ [384, 1536], depending on the model.
Each dimension doesn’t have an interpretable meaning — but collectively, they encode semantic context.
For example, the phrase “cloud storage” might produce:
$$ v = [0.21, 0.47, -0.05, ..., 0.32] $$
A similar phrase like “file hosting service” yields a nearby vector:
$$ u = [0.19, 0.50, -0.07, ..., 0.30] $$
Their cosine similarity:
$$ \text{sim}(u, v) = \frac{u \cdot v}{|u||v|} \approx 0.98 $$
That geometric closeness means they’re neighbors in semantic space — the essence of dense retrieval.
4. Common Embedding Models
| Provider | Model | Strengths | Use Case |
|---|---|---|---|
| Hugging Face (SentenceTransformers) | all-MiniLM-L6-v2 |
Fast, open, good baseline | General-purpose |
| Cohere | embed-english-v3 |
Strong multilingual and topic clustering | Enterprise search |
| OpenAI | text-embedding-3-large |
High-quality, robust | Precision-critical RAG |
| Mistral | mistral-embed |
Efficient, small footprint | Lightweight pipelines |
| SageMaker-hosted | Any | Fully managed, scalable | Production workloads |
Each model has its own embedding space, so embeddings are not interchangeable across providers.
That’s why modularity (configurable EMBEDDING_PROVIDER and EMBEDDING_MODEL) is crucial in the Default RAG Stack.
5. Measuring Similarity
Once all chunks and queries are embedded, retrieval comes down to comparing their vectors.
The three main distance metrics are:
- Cosine similarity:
$$ \text{cosine}(a,b) = \frac{a \cdot b}{|a||b|} $$- Most common; scale-independent.
- Dot product:
$$ \text{dot}(a,b) = a \cdot b $$- Simpler, used when vectors are normalized.
- Euclidean distance:
$$ d(a,b) = |a - b|_2 $$- Intuitive geometric distance, but less robust for high dimensions.
Vector databases use these metrics to compute “nearest neighbors” efficiently — finding which stored chunks are most relevant to a query vector.
6. Enter the Vector Database
A vector database is a specialized data system for storing and searching embeddings at scale.
In RAG, it acts as the semantic memory — the long-term knowledge store that the retriever consults.
Key responsibilities:
- Store vectors + metadata
- Perform Approximate Nearest Neighbor (ANN) search
- Support filtering and pagination
- Scale efficiently as collections grow
7. Qdrant: The Engine of Semantic Recall
Qdrant is the chosen vector store in the Default RAG Stack because it’s open-source, production-grade, and integrates cleanly with LangChain.
Core concepts:
- Collection: like a database table; each stores vectors and payloads.
- Vector: dense numerical array (e.g., 768D).
- Payload: key-value metadata for filtering.
- Index: built using HNSW (Hierarchical Navigable Small World graph).
HNSW enables sublinear-time similarity search by connecting each vector to its closest neighbors in a graph structure. Search then traverses the graph efficiently — much faster than comparing all pairs.
Simplified workflow:
- $$ \text{insert}(v_i, \text{payload}_i) \rightarrow \text{HNSW Graph Update} $$
- $$ \text{query}(q) \rightarrow \text{nearest neighbors via greedy + beam search} $$
8. Storing Knowledge Efficiently
Each Qdrant record contains:
{
"id": "chunk_00042",
"vector": [...],
"payload": {
"source": "Earnings_Call_Q3.pdf",
"page": 7,
"created_at": "2025-10-12"
}
}
Qdrant supports filtering by any payload key, so queries like
“find relevant paragraphs from 2025 earnings calls only”
can be executed directly at retrieval time:
results = qdrant.search(
query_vector=query_vec,
limit=5,
filter={"must": [{"key": "source", "match": {"value": "Earnings"}}]}
)
This turns metadata into first-class retrieval logic.
9. Vector Indexing: The Hidden Hero
Indexing determines both speed and recall quality.
Qdrant’s HNSW structure optimizes the balance between search precision and latency using graph-based traversal.
Rough intuition:
- Build a graph where each vector connects to its MMM nearest neighbors.
- Search navigates the graph greedily, keeping a beam of candidate nodes.
- The larger MMM and beam width, the higher the recall but the slower the search.
Trade-off tuning (in Qdrant config):
| Parameter | Meaning | Effect |
|---|---|---|
ef_construct |
# of neighbors to consider during build | Higher = better recall, more memory |
ef_search |
# of candidates checked during query | Higher = better recall, slower query |
M |
Graph degree | Controls connectivity density |
10. Vector Databases vs Traditional Databases
| Aspect | Vector DB (Qdrant) | Relational DB (PostgreSQL) |
|---|---|---|
| Query Type | Nearest neighbor (semantic) | Exact match or join |
| Data | High-dimensional vectors | Structured rows/columns |
| Storage | Optimized for float arrays | Optimized for integers/strings |
| Retrieval Time | O(logn) (approximate) | O(n) (exact) |
| Use Case | Semantic search, recommendations | Transactions, analytics |
They serve different purposes — RAG uses both: Qdrant for meaning, Django/PostgreSQL for metadata persistence.
11. Visual Analogy: The Semantic Galaxy
Imagine each chunk of text as a star in a galaxy.
Similar ideas cluster together into constellations — “finance,” “product launches,” “customer feedback.”
The vector database is the telescope that, given a query star, instantly zooms into its neighboring constellation.
Embeddings determine where stars are placed; Qdrant determines how to find them quickly.
12. Key Takeaways
- Embeddings encode meaning into math.
- Semantic similarity enables understanding beyond exact words.
- Qdrant efficiently stores and retrieves those vectors with metadata filters.
- Graph-based indexing (HNSW) balances speed and recall.
- Together, embeddings + vector databases form the semantic core of every RAG system.
Next in the series → Part 4: Retrieval & Generation Pipeline (LangChain Core)