Scaling Embeddings to a Billion Vectors: Architecture Lessons
Back to all articles
AI Engineering
21 min read8 min read

Scaling Embeddings to a Billion Vectors: Architecture Lessons

How we scaled our embedding system from millions to billions of vectors. Covers partitioning, quantization, caching, and index optimization.

Debasish Maji
Debasish Maji
AI Engineering Lead
April 12, 2026
EmbeddingsScalingVector DBInfrastructureProduction

The Scaling Challenge

At 1 million vectors, everything works. Your vector database responds in milliseconds. Indexing is fast. Life is good.

At 100 million vectors, cracks appear. Query latency creeps up. Index builds take hours. Memory usage becomes a concern.

At 1 billion vectors, everything you thought you knew breaks. This is the story of how we scaled our embedding infrastructure 1000x while keeping p99 latency under 50ms.

•••

The Starting Point

Our initial architecture was simple:

ComponentChoiceCapacity
Vector DBSingle Pinecone pod1M vectors
Embedding modelOpenAI ada-0021536 dimensions
Index typeHNSWDefault settings
Query patternSingle vector similarityTop-10 results
This worked perfectly until it did not.

•••

Scaling Phase 1: Vertical Scaling (1M to 10M)

The first instinct: throw more resources at it.

What We Tried

ResourceBeforeAfterImpact
Memory16GB64GB4x capacity
CPU cores4162x query throughput
StorageSSDNVMe3x index build speed

Results

Metric1M vectors10M vectors
p50 latency8ms12ms
p99 latency25ms45ms
Index build5 min2 hours
Monthly cost$200$800
Vertical scaling got us 10x with acceptable performance. But the next 10x would not be so easy.

•••

Scaling Phase 2: Horizontal Sharding (10M to 100M)

Single-node limits forced us to distribute.

Sharding Strategies

We evaluated three approaches:

StrategyProsCons
Random shardingSimple, even distributionQuery all shards
Hash-basedDeterministic routingQuery all shards
Semantic clusteringQuery subset of shardsComplex rebalancing
We chose semantic clustering - grouping similar vectors together.

Semantic Clustering Implementation

The approach clusters vectors by content type, routes queries to relevant shards, and only fans out when necessary.

ClusterContent TypeSize
0Technical docs25M
1Product content30M
2Support articles20M
3User content15M
4General10M

Query Routing

For each query, we first determine likely clusters, then query only those shards:

Query TypeShards QueriedLatency Savings
Technical question1-2 shards60% faster
Product search1-2 shards55% faster
General search3-5 shardsNo savings
Average queries hit 2.3 shards instead of all 5.

Phase 2 Results

Metric10M vectors100M vectors
p50 latency12ms18ms
p99 latency45ms52ms
Query throughput500 qps2000 qps
Monthly cost$800$3,500
•••

Scaling Phase 3: The Billion Vector Architecture (100M to 1B)

The jump to a billion required rethinking everything.

Architecture Overview

The final architecture has three tiers:

TierPurposeTechnology
HotFrequent queriesIn-memory HNSW
WarmRecent contentSSD-backed index
ColdArchiveDisk-based with on-demand loading

Tiered Storage Strategy

TierVectorsQuery LatencyStorage Cost
Hot50M15msHigh
Warm300M35msMedium
Cold650M150msLow
85% of queries hit hot tier data. Only 2% need cold tier access.

Index Optimization

At billion scale, index configuration is critical:

ParameterSmall ScaleBillion ScaleWhy
M (connections)1632Better recall at scale
efConstruction200400Quality over build time
efSearch100150Balance recall/latency

Quantization for Memory Efficiency

Full 1536-dimension float32 vectors use 6KB each. At 1 billion vectors, that is 6TB just for vectors.

QuantizationMemory per VectorRecall Impact
None (float32)6144 bytesBaseline
float163072 bytes-0.1%
int81536 bytes-0.5%
Product Quantization384 bytes-2%
We use int8 for hot tier (minimal recall loss) and PQ for cold tier (acceptable for archive).

Memory Calculation

TierVectorsQuantizationMemory
Hot50Mint877GB
Warm300Mint8461GB
Cold650MPQ250GB
Total1BMixed788GB
Compared to 6TB without quantization - 87% memory reduction.

•••

Query Path Optimization

Every millisecond matters at scale.

Query Pipeline

StageTime BudgetOptimization
Embedding15msBatch, cache popular
Routing2msPrecomputed cluster model
Index search25msTiered, parallel shards
Reranking5msGPU-accelerated
Response3msPrecomputed metadata
Total budget: 50ms p99.

Caching Layers

CacheHit RateLatency Saved
Query embedding15%15ms
Full result8%45ms
Metadata40%3ms
Even modest cache hit rates significantly improve p99.

Parallel Shard Queries

When querying multiple shards, we query in parallel:

Approach3 Shards5 Shards
Sequential75ms125ms
Parallel28ms32ms
Parallel queries are essential for multi-shard searches.

•••

Infrastructure Decisions

Vector Database Selection

We evaluated options for billion-scale:

DatabaseMax ScaleManagedCost
Pinecone1B+Yes$$$$
Weaviate100M+Self/Managed$$
Qdrant1B+Self/Managed$$
Milvus1B+Self-hosted$
pgvector10MSelf-hosted$
We chose Qdrant for cost-efficiency and control, with Pinecone as a fallback for critical workloads.

Infrastructure Costs at Scale

Component100M1BNotes
Vector storage$3,500$12,000Tiered reduces cost
Compute$2,000$8,000Auto-scaling
Embedding generation$1,500$5,000Batching helps
Total monthly$7,000$25,000Linear-ish scaling
•••

Operational Lessons

Index Rebuilds

At billion scale, full rebuilds are expensive:

Vector CountFull Rebuild Time
1M5 minutes
100M8 hours
1B3 days
We use incremental updates instead:

OperationTimeImpact
Single insert50msNone
Batch insert (1000)2sNone
Index optimization30 minBrief latency spike
Full rebuild3 daysRequires replica

Monitoring at Scale

MetricWarningCritical
p99 latency75ms100ms
Cache hit rateBelow 10%Below 5%
Shard imbalance20% deviation40% deviation
Memory usage80%90%
Index fragmentation30%50%
•••

Results: Before and After

Metric1M (Start)1B (Final)
Vector count1M1,000M
p50 latency8ms22ms
p99 latency25ms48ms
Query throughput200 qps5,000 qps
Monthly cost$200$25,000
Cost per 1M vectors$200$25
1000x scale with only 3x latency increase and 8x reduction in per-vector cost.

•••

Key Takeaways

  1. 1Vertical scaling has limits - Useful for 10x, but horizontal is required beyond that.
  1. 2Semantic sharding beats random - Query fewer shards by clustering similar content.
  1. 3Tiered storage is essential - Not all vectors need fast access. Tier by access patterns.
  1. 4Quantization is your friend - 87% memory reduction with minimal recall loss.
  1. 5Cache aggressively - Query embeddings, results, and metadata all benefit from caching.
  1. 6Parallel everything - Sequential shard queries kill latency at scale.
  1. 7Plan for incremental updates - Full rebuilds become impractical. Build for incremental from day one.

Billion-scale vector search is achievable with careful architecture. The key is accepting that the simple single-node approach will not work and designing for distribution from the start.

Found this helpful?

Share it with others who might benefit

TweetShare

Related articles