Embeddings are the foundation of semantic search, RAG systems, recommendation engines, and more. They transform text, images, and other data into dense vectors that capture meaning.
Choosing the right embedding model and using it correctly can make or break your AI application. This guide covers everything you need to know.
| Property | Description | Example |
|---|
| Semantic similarity | Meaning closeness | "car" near "automobile" |
| Relationships | Analogies | king - man + woman = queen |
| Context | Surrounding meaning | "bank" differs by context |
| Domain knowledge | Specialized concepts | Medical terms clustered |
| Dimension | Storage | Speed | Quality |
|---|
| 384 | Small | Fast | Good |
| 768 | Medium | Medium | Better |
| 1024 | Large | Slower | Best |
| 1536 | Very large | Slow | Marginal gains |
| 3072 | Huge | Very slow | Diminishing returns |
| Use Case | Recommended Dimension | Reasoning |
|---|
| Real-time search | 384-768 | Speed priority |
| High accuracy RAG | 1024-1536 | Quality priority |
| Large scale (1B+ vectors) | 384-512 | Storage critical |
| Specialized domain | 768-1024 | Balance |
| Model | Dimensions | Context | Performance | Cost |
|---|
| OpenAI text-embedding-3-large | 3072 | 8191 | Excellent | $$$ |
| OpenAI text-embedding-3-small | 1536 | 8191 | Very good | $$ |
| Cohere embed-v3 | 1024 | 512 | Excellent | $$ |
| Voyage AI | 1024 | 16000 | Excellent | $$ |
| BGE-large | 1024 | 512 | Very good | Free |
| E5-large-v2 | 1024 | 512 | Very good | Free |
| all-MiniLM-L6 | 384 | 256 | Good | Free |
| Model | Retrieval | Classification | Clustering | Average |
|---|
| text-embedding-3-large | 64.2 | 75.1 | 49.8 | 63.0 |
| Cohere embed-v3 | 63.8 | 74.5 | 50.1 | 62.8 |
| BGE-large | 62.1 | 73.2 | 48.9 | 61.4 |
| E5-large-v2 | 61.8 | 72.8 | 48.2 | 60.9 |
| all-MiniLM-L6 | 56.2 | 68.1 | 44.3 | 56.2 |
| Model | Cost per 1M tokens | 1M Documents (500 tokens avg) |
|---|
| text-embedding-3-large | $0.13 | $65 |
| text-embedding-3-small | $0.02 | $10 |
| Cohere embed-v3 | $0.10 | $50 |
| Self-hosted BGE | ~$0.01 | ~$5 |
| Chunk Size | Retrieval Precision | Context Relevance | Cost |
|---|
| 128 tokens | High | Low (fragments) | High |
| 256 tokens | Good | Medium | Medium |
| 512 tokens | Medium | Good | Low |
| 1024 tokens | Lower | High | Very low |
| Method | Description | Best For |
|---|
| Fixed size | Split at token count | Simple, consistent |
| Sentence | Split at sentence boundaries | Natural breaks |
| Paragraph | Split at paragraphs | Coherent chunks |
| Semantic | Split by meaning | Best quality |
| Recursive | Hierarchical splitting | Documents with structure |
| Overlap | Benefit | Trade-off |
|---|
| 0% | Minimum storage | Context loss at boundaries |
| 10% | Some continuity | Slight redundancy |
| 20% | Good continuity | More storage |
| 50% | Maximum continuity | 2x storage |
| Document Type | Chunk Size | Overlap | Method |
|---|
| Technical docs | 512 | 20% | Recursive |
| Articles | 256-512 | 10% | Paragraph |
| Code | 256 | 10% | Semantic |
| Conversations | 128-256 | 20% | Message-based |
| Legal documents | 512-1024 | 15% | Section-based |
| Index Type | Build Time | Query Time | Recall | Memory |
|---|
| Flat (exact) | O(n) | O(n) | 100% | Low |
| IVF | O(n) | O(sqrt(n)) | 95-99% | Medium |
| HNSW | O(n log n) | O(log n) | 98-99.5% | High |
| PQ | O(n) | O(n/compression) | 90-95% | Very low |
| IVF-PQ | O(n) | O(sqrt(n)/compression) | 85-95% | Low |
| Scale | Latency Need | Recall Need | Recommended |
|---|
| Under 100K | Any | Any | Flat |
| 100K-1M | Low | High | HNSW |
| 1M-100M | Medium | High | IVF + HNSW |
| 100M+ | Medium | Medium | IVF-PQ |
| 1B+ | Any | Medium | Distributed IVF-PQ |
| Parameter | Low Value | High Value | Trade-off |
|---|
| M | 8 | 64 | Memory vs recall |
| ef_construction | 64 | 512 | Build time vs recall |
| ef_search | 32 | 256 | Query time vs recall |
| Use Case | M | ef_construction | ef_search |
|---|
| Speed priority | 16 | 100 | 50 |
| Balanced | 32 | 200 | 100 |
| Recall priority | 48 | 400 | 200 |
| Scenario | Fine-Tune? | Why |
|---|
| Domain-specific vocabulary | Yes | Better representation |
| Poor out-of-box performance | Yes | Improve relevance |
| Unique similarity definition | Yes | Custom distance |
| General improvement | Maybe | Expensive, risky |
| Small dataset | No | Overfitting risk |
| Data Size | Expected Improvement | Risk |
|---|
| Under 1K pairs | 0-5% | High overfitting |
| 1K-10K pairs | 5-15% | Medium overfitting |
| 10K-100K pairs | 10-25% | Low risk |
| 100K+ pairs | 15-30% | Minimal risk |
| Domain | Base Model | Fine-Tuned | Improvement |
|---|
| Legal | 72% | 89% | +24% |
| Medical | 68% | 85% | +25% |
| E-commerce | 75% | 88% | +17% |
| Technical | 78% | 91% | +17% |
| Stage | Action | Latency |
|---|
| Preprocessing | Clean, normalize text | 5ms |
| Chunking | Split into segments | 10ms |
| Embedding | Generate vectors | 50-200ms |
| Indexing | Add to vector store | 5-20ms |
| Total | End-to-end | 70-235ms |
| Batch Size | Throughput | Latency per Item |
|---|
| 1 | Baseline | Baseline |
| 8 | 6x | 15% of baseline |
| 32 | 20x | 5% of baseline |
| 128 | 50x | 2% of baseline |
| Strategy | Hit Rate | Storage | Freshness |
|---|
| No cache | 0% | 0 | Always fresh |
| Document hash | 80-95% | Medium | Stale risk |
| Content hash | 95%+ | Medium | Always fresh |
| TTL-based | 70-90% | Low | Configurable |
| Pitfall | Problem | Solution |
|---|
| Wrong chunk size | Poor retrieval | Test multiple sizes |
| No preprocessing | Noise in embeddings | Clean text first |
| Ignoring context | Missing nuance | Use longer chunks or context |
| Single model | No fallback | Multi-model strategy |
| No monitoring | Quality drift | Track retrieval metrics |
| Metric | Description | Target |
|---|
| Recall@K | Relevant in top K | Over 90% |
| Precision@K | Relevant of top K | Over 70% |
| MRR | Reciprocal rank | Over 0.8 |
| NDCG | Ranked relevance | Over 0.85 |
| Metric | Description | Target |
|---|
| Embedding latency | Time to embed | Under 100ms |
| Query latency | Time to search | Under 50ms |
| Throughput | Queries per second | Over 100 QPS |
| Index size | Storage required | Predictable |
- 1Dimension is a trade-off - Higher is not always better. Match dimension to your scale and latency needs.
- 2Chunking matters enormously - The right chunk size and overlap can improve retrieval by 20-30%.
- 3HNSW for most use cases - Until you hit billions of vectors, HNSW provides the best recall/speed balance.
- 4Fine-tuning is powerful for domains - 15-25% improvement is achievable with domain-specific training.
- 5Batch for efficiency - Batching embeddings can improve throughput by 50x.
- 6Monitor retrieval quality - Embeddings drift. Track Recall@K and MRR continuously.
Embeddings are the silent foundation of modern AI systems. Getting them right is the difference between a system that works and a system that delights.