Thrive AI Editorial
Thrive AI Editorial · AI-assisted worked example, not a member conversation.
2026-10-07
First define what must generalize
If production predictions concern unseen customers, evaluation should hold out customer identities, not just rows. Shared identities let a model exploit stable customer-specific signals that will be unavailable for a new customer. A high row-split score can therefore answer the wrong question.
An overlapping-customer split is not automatically wrong for every task. If you predict future activity for existing customers, a time-respecting split may be appropriate. In that case, every feature must still be available at the prediction timestamp; future labels, aggregates and lookups must not enter training.
Target: unseen customers
customer IDs -> disjoint groups -> fit on training group only
-> evaluate on held-out group
See the mismatch without training a complex model
Python 3, standard library only. A lookup memorizes customer labels. The example is intentionally constructed, so its percentages illustrate the failure mode rather than estimate real model quality.
rows = [
{"customer": customer, "repeat": repeat, "label": customer % 2}
for customer in range(8)
for repeat in range(2)
]
def fit(train):
return {row["customer"]: row["label"] for row in train}
def accuracy(model, test):
predictions = [
model.get(row["customer"], 0) == row["label"]
for row in test
]
return sum(predictions) / len(predictions)
row_train = [row for row in rows if row["repeat"] == 0]
row_test = [row for row in rows if row["repeat"] == 1]
group_train = [row for row in rows if row["customer"] < 4]
group_test = [row for row in rows if row["customer"] >= 4]
overlap = {r["customer"] for r in row_train} & {
r["customer"] for r in row_test
}
assert len(overlap) == 8
assert not (
{r["customer"] for r in group_train}
& {r["customer"] for r in group_test}
)
assert accuracy(fit(row_train), row_test) == 1.0
assert accuracy(fit(group_train), group_test) == 0.5
print("row split:", accuracy(fit(row_train), row_test))
print("group split:", accuracy(fit(group_train), group_test))
Expected output:
row split: 1.0
group split: 0.5
The row example deterministically alternates repeated records rather than drawing a random split, to make the overlap easy to inspect. Random row splitting can create the same overlap, but its exact amount depends on sampling.
Fix more than the split
- Decide whether the holdout unit is a customer, patient, document, organization, session or time period. Record that deployment assumption before tuning.
- Make the groups disjoint for unseen-entity evaluation. For repeated experiments, use group-aware cross-validation and inspect class balance across folds.
- Learn imputers, scalers, feature selection and encoders on each training fold only. A pipeline prevents many preprocessing mistakes, but it cannot repair leaked features or the wrong split strategy.
- Keep the final test set out of prompt, feature and model selection. Tuning repeatedly on it makes it part of development.
- Inspect duplicates and near-duplicates across partitions. Changing a record identifier does not make copied content independent.
For retrieval-augmented generation (RAG), the analogous question is whether near-identical questions, document versions or answer-bearing artifacts cross your chosen evaluation boundary. This does not mean all retrieval documents must be excluded: production RAG is supposed to retrieve useful knowledge. Define what information should be available and what must remain unseen for the task you are measuring.
Evidence to include in a follow-up question
Share group counts per split, overlap checks, timestamps, the feature-availability rule and how preprocessing was fitted. An unexplained score alone cannot establish leakage or prove that it has been eliminated.
Source checked 2026-10-07: scikit-learn: Common pitfalls and data leakage. Its key guidance is to split before fitting preprocessing and never learn transformations from the test set. The synthetic entity-overlap demonstration and RAG analogy are our AI-assisted explanation, not benchmark results.
Sign in to vote or contribute. Scores count actual member votes, not editorial endorsements.