Thrive AI Editorial
Thrive AI Editorial · AI-assisted worked example, not a member conversation.
2026-10-07
Filtering and truncation do not commute
Taking the global top-k and then removing disallowed results can leave fewer than k results. Allowed documents just below the original cutoff never get a chance to enter the final list. Filtering the eligible candidate set before selecting its top-k solves this particular truncation problem.
Real approximate nearest-neighbor engines have additional tradeoffs involving shards, graph traversal, selectivity and recall. Microsoft documents distinct prefilter, postfilter and strict-postfilter modes. Our example corresponds to the simple global-cutoff failure; it does not reproduce every engine's postfilter implementation.
global ranking -> truncate -> filter -> possibly empty context
eligible candidates -> rank -> truncate -> eligible top-k
A minimal counterexample
Python 3, standard library only. Scores are fixed synthetic relevance scores, not embedding measurements.
ranked = [
{"id": "x1", "collection": "other", "score": 0.99},
{"id": "x2", "collection": "other", "score": 0.98},
{"id": "a1", "collection": "allowed", "score": 0.97},
{"id": "a2", "collection": "allowed", "score": 0.96},
]
k = 2
eligible = lambda row: row["collection"] == "allowed"
postfiltered = [row["id"] for row in ranked[:k] if eligible(row)]
prefiltered = [row["id"] for row in ranked if eligible(row)][:k]
assert postfiltered == []
assert prefiltered == ["a1", "a2"]
assert all(eligible(row) for row in ranked if row["id"] in prefiltered)
print("after cutoff:", postfiltered)
print("before cutoff:", prefiltered)
Expected output:
after cutoff: []
before cutoff: ['a1', 'a2']
Diagnose your retrieval pipeline
Log safe identifiers and counts at each stage: eligible corpus, retrieved candidates, filtered candidates, reranked candidates and final context. Keep raw private document text out of broadly accessible traces. Compare candidate recall before evaluating generated answers.
Increasing the initial candidate pool can reduce this failure but does not guarantee sufficient eligible matches at every selectivity. Where supported, push eligibility constraints into retrieval. Measure latency and recall with representative sparse and dense permission sets rather than extrapolating from one collection.
Metadata filtering is not automatically complete authorization. Resolve permissions from trusted application state; do not let the model or an arbitrary client choose a broader tenant filter. Recheck access before exposing documents, and invalidate stale permissions and caches. Keeping an unauthorized document out of the answer is too late if it already reached an unauthorized model context or log.
Verification and limitations
Test no eligible documents, exactly k eligible documents, eligible documents below the global cutoff, and permissions changing between retrieval and display. Check zero forbidden documents in the model context as a separate invariant from retrieval quality.
Our fixed ordered list has exact ranking and no sharding. A production approximate index may still miss relevant eligible candidates. Diagnose that separately with an exact-search baseline on a manageable sample. Do not present this toy output as measured vector-database performance.
Source checked 2026-10-08: Azure AI Search vector query filters. Filtering-mode and recall tradeoffs are documented there. The ranking counterexample and authorization discussion are original AI-assisted teaching material; no Azure service was called.
Sign in to vote or contribute. Scores count actual member votes, not editorial endorsements.