At 1 million vectors, everything works. Your vector database responds in milliseconds. Indexing is fast. Life is good.
At 100 million vectors, cracks appear. Query latency creeps up. Index builds take hours. Memory usage becomes a concern.
At 1 billion vectors, everything you thought you knew breaks. This is the story of how we scaled our embedding infrastructure 1000x while keeping p99 latency under 50ms.
Our initial architecture was simple:
| Component | Choice | Capacity |
|---|
| Vector DB | Single Pinecone pod | 1M vectors |
| Embedding model | OpenAI ada-002 | 1536 dimensions |
| Index type | HNSW | Default settings |
| Query pattern | Single vector similarity | Top-10 results |
This worked perfectly until it did not.The first instinct: throw more resources at it.
| Resource | Before | After | Impact |
|---|
| Memory | 16GB | 64GB | 4x capacity |
| CPU cores | 4 | 16 | 2x query throughput |
| Storage | SSD | NVMe | 3x index build speed |
| Metric | 1M vectors | 10M vectors |
|---|
| p50 latency | 8ms | 12ms |
| p99 latency | 25ms | 45ms |
| Index build | 5 min | 2 hours |
| Monthly cost | $200 | $800 |
Vertical scaling got us 10x with acceptable performance. But the next 10x would not be so easy.Single-node limits forced us to distribute.
We evaluated three approaches:
| Strategy | Pros | Cons |
|---|
| Random sharding | Simple, even distribution | Query all shards |
| Hash-based | Deterministic routing | Query all shards |
| Semantic clustering | Query subset of shards | Complex rebalancing |
We chose semantic clustering - grouping similar vectors together.The approach clusters vectors by content type, routes queries to relevant shards, and only fans out when necessary.
| Cluster | Content Type | Size |
|---|
| 0 | Technical docs | 25M |
| 1 | Product content | 30M |
| 2 | Support articles | 20M |
| 3 | User content | 15M |
| 4 | General | 10M |
For each query, we first determine likely clusters, then query only those shards:
| Query Type | Shards Queried | Latency Savings |
|---|
| Technical question | 1-2 shards | 60% faster |
| Product search | 1-2 shards | 55% faster |
| General search | 3-5 shards | No savings |
Average queries hit 2.3 shards instead of all 5.| Metric | 10M vectors | 100M vectors |
|---|
| p50 latency | 12ms | 18ms |
| p99 latency | 45ms | 52ms |
| Query throughput | 500 qps | 2000 qps |
| Monthly cost | $800 | $3,500 |
The jump to a billion required rethinking everything.
The final architecture has three tiers:
| Tier | Purpose | Technology |
|---|
| Hot | Frequent queries | In-memory HNSW |
| Warm | Recent content | SSD-backed index |
| Cold | Archive | Disk-based with on-demand loading |
| Tier | Vectors | Query Latency | Storage Cost |
|---|
| Hot | 50M | 15ms | High |
| Warm | 300M | 35ms | Medium |
| Cold | 650M | 150ms | Low |
85% of queries hit hot tier data. Only 2% need cold tier access.At billion scale, index configuration is critical:
| Parameter | Small Scale | Billion Scale | Why |
|---|
| M (connections) | 16 | 32 | Better recall at scale |
| efConstruction | 200 | 400 | Quality over build time |
| efSearch | 100 | 150 | Balance recall/latency |
Full 1536-dimension float32 vectors use 6KB each. At 1 billion vectors, that is 6TB just for vectors.
| Quantization | Memory per Vector | Recall Impact |
|---|
| None (float32) | 6144 bytes | Baseline |
| float16 | 3072 bytes | -0.1% |
| int8 | 1536 bytes | -0.5% |
| Product Quantization | 384 bytes | -2% |
We use int8 for hot tier (minimal recall loss) and PQ for cold tier (acceptable for archive).| Tier | Vectors | Quantization | Memory |
|---|
| Hot | 50M | int8 | 77GB |
| Warm | 300M | int8 | 461GB |
| Cold | 650M | PQ | 250GB |
| Total | 1B | Mixed | 788GB |
Compared to 6TB without quantization - 87% memory reduction.Every millisecond matters at scale.
| Stage | Time Budget | Optimization |
|---|
| Embedding | 15ms | Batch, cache popular |
| Routing | 2ms | Precomputed cluster model |
| Index search | 25ms | Tiered, parallel shards |
| Reranking | 5ms | GPU-accelerated |
| Response | 3ms | Precomputed metadata |
Total budget: 50ms p99.| Cache | Hit Rate | Latency Saved |
|---|
| Query embedding | 15% | 15ms |
| Full result | 8% | 45ms |
| Metadata | 40% | 3ms |
Even modest cache hit rates significantly improve p99.When querying multiple shards, we query in parallel:
| Approach | 3 Shards | 5 Shards |
|---|
| Sequential | 75ms | 125ms |
| Parallel | 28ms | 32ms |
Parallel queries are essential for multi-shard searches.We evaluated options for billion-scale:
| Database | Max Scale | Managed | Cost |
|---|
| Pinecone | 1B+ | Yes | $$$$ |
| Weaviate | 100M+ | Self/Managed | $$ |
| Qdrant | 1B+ | Self/Managed | $$ |
| Milvus | 1B+ | Self-hosted | $ |
| pgvector | 10M | Self-hosted | $ |
We chose Qdrant for cost-efficiency and control, with Pinecone as a fallback for critical workloads.| Component | 100M | 1B | Notes |
|---|
| Vector storage | $3,500 | $12,000 | Tiered reduces cost |
| Compute | $2,000 | $8,000 | Auto-scaling |
| Embedding generation | $1,500 | $5,000 | Batching helps |
| Total monthly | $7,000 | $25,000 | Linear-ish scaling |
At billion scale, full rebuilds are expensive:
| Vector Count | Full Rebuild Time |
|---|
| 1M | 5 minutes |
| 100M | 8 hours |
| 1B | 3 days |
We use incremental updates instead:| Operation | Time | Impact |
|---|
| Single insert | 50ms | None |
| Batch insert (1000) | 2s | None |
| Index optimization | 30 min | Brief latency spike |
| Full rebuild | 3 days | Requires replica |
| Metric | Warning | Critical |
|---|
| p99 latency | 75ms | 100ms |
| Cache hit rate | Below 10% | Below 5% |
| Shard imbalance | 20% deviation | 40% deviation |
| Memory usage | 80% | 90% |
| Index fragmentation | 30% | 50% |
| Metric | 1M (Start) | 1B (Final) |
|---|
| Vector count | 1M | 1,000M |
| p50 latency | 8ms | 22ms |
| p99 latency | 25ms | 48ms |
| Query throughput | 200 qps | 5,000 qps |
| Monthly cost | $200 | $25,000 |
| Cost per 1M vectors | $200 | $25 |
1000x scale with only 3x latency increase and 8x reduction in per-vector cost.- 1Vertical scaling has limits - Useful for 10x, but horizontal is required beyond that.
- 2Semantic sharding beats random - Query fewer shards by clustering similar content.
- 3Tiered storage is essential - Not all vectors need fast access. Tier by access patterns.
- 4Quantization is your friend - 87% memory reduction with minimal recall loss.
- 5Cache aggressively - Query embeddings, results, and metadata all benefit from caching.
- 6Parallel everything - Sequential shard queries kill latency at scale.
- 7Plan for incremental updates - Full rebuilds become impractical. Build for incremental from day one.
Billion-scale vector search is achievable with careful architecture. The key is accepting that the simple single-node approach will not work and designing for distribution from the start.