Thrive AI Engineering
← All questions

rag · Editorial starter

Why should I not add keyword and vector scores directly in hybrid retrieval?

By Thrive AI Editorial · Created 2026-10-07 · Updated 2026-10-07

Thrive AI Editorial · AI-assisted worked example, not a member conversation.

Keyword retrieval returns scores around twenty while vector similarity returns values below one. Adding them makes the keyword ranking dominate. How can the two ranked lists be combined without assuming their raw scores share a scale?

This is an editorial ranking example with synthetic document identifiers, not a retrieval benchmark or claim that one fusion method always wins.

0 score

Sign in to vote or contribute. Scores count actual member votes, not editorial endorsements.

1 published answer

A selected solution is chosen by the question author or moderator, not an independent certification. Check assumptions before using any code in production.

Thrive AI Editorial

Thrive AI Editorial · AI-assisted worked example, not a member conversation.

2026-10-07

Scores from different rankers are not interchangeable units

A keyword ranking score and a vector similarity are produced by different functions. Adding them with equal numeric weights does not give each method equal influence. Normalization can help, but it introduces assumptions that need evaluation.

Reciprocal Rank Fusion (RRF) is an alternative that uses positions within each ranked list rather than raw score magnitudes. For each document, sum 1 / (constant + rank) across the lists in which it appears. Documents ranked near the top by several retrievers accumulate support.

The fusion constant is not the number of nearest neighbors requested from a vector index. Calling both values k can conceal a configuration error.

Code / output
keyword ranked IDs ----\
                        -> rank-based fusion -> candidate list -> evaluate
vector ranked IDs -----/

Compute the fused order

Python 3, standard library only. The rankings and constant are illustrative. Each input list contains a document at most once.

Code / output
from collections import defaultdict

rankings = [["A", "B", "C"], ["B", "D", "A"]]
fusion_constant = 60
scores = defaultdict(float)
for ranking in rankings:
    assert len(ranking) == len(set(ranking))
    for rank, document in enumerate(ranking, start=1):
        scores[document] += 1 / (fusion_constant + rank)

ordered = sorted(scores, key=lambda doc: (-scores[doc], doc))
assert ordered == ["B", "A", "D", "C"]
assert scores["B"] > scores["A"] > scores["D"] > scores["C"]
print("keyword:", rankings[0])
print("vector:", rankings[1])
print("fused:", ordered)

Expected output:

Code / output
keyword: ['A', 'B', 'C']
vector: ['B', 'D', 'A']
fused: ['B', 'A', 'D', 'C']

What this fixes and what it cannot fix

RRF removes dependence on the raw magnitudes in this combination step. It does not establish that any retrieved document is relevant. If both lists favor the same irrelevant item, fusion can reinforce the mistake.

Candidate depth matters. A relevant document absent from every input list cannot be recovered by a rank-fusion formula. Deduplicate stable document identifiers within each list; accidentally counting repeated chunks as repeated support changes the scoring rule.

Apply eligibility constraints before material is exposed, including before reranking with an external model. Fuse comparable retrieval units deliberately: mixing document IDs, chunk IDs and different document versions without a mapping creates misleading duplicates.

Verify against the actual retrieval task

Compare keyword-only, vector-only and fused candidates on labeled queries. Inspect recall at the chosen context size and the quality of downstream answers. Segment by exact identifiers, rare terminology, paraphrases and queries needing multiple pieces of evidence.

Do not choose a constant because the demonstration uses sixty. Treat candidate depth, fusion weights and optional reranking as evaluated configuration. Keep an untouched test set, record the index snapshot and report latency alongside quality. An improvement on synthetic rankings is not evidence of a production improvement.

Source checked 2026-10-08: Microsoft hybrid-search ranking and RRF. The reciprocal-rank formula and distinction between ranking scores are documented there. The worked ranking and assertions are original AI-assisted material; no search service was queried.

0 score

Sign in to vote or contribute. Scores count actual member votes, not editorial endorsements.

Verify your sign-in to contribute an answer

Reading is free. Posting requires a verified account and moderation; no course purchase is required.