Thrive AI Engineering
← All questions

machine-learning · Editorial starter

Why can a random train/test split look strong but fail on new customers?

By Thrive AI Editorial · Created 2026-10-07 · Updated 2026-10-07

Thrive AI Editorial · AI-assisted worked example, not a member conversation.

A dataset has several rows per customer. A row-level random split places records for the same customer in both training and evaluation. A model performs well in that test but poorly for customers it has never seen. Is the split measuring the intended deployment task, and what should change?

This editorial example uses synthetic identifiers and a simple memorization baseline to isolate entity overlap. It does not claim a measured result on a real business dataset.

0 score

Sign in to vote or contribute. Scores count actual member votes, not editorial endorsements.

1 published answer

A selected solution is chosen by the question author or moderator, not an independent certification. Check assumptions before using any code in production.

Thrive AI Editorial

Thrive AI Editorial · AI-assisted worked example, not a member conversation.

2026-10-07

First define what must generalize

If production predictions concern unseen customers, evaluation should hold out customer identities, not just rows. Shared identities let a model exploit stable customer-specific signals that will be unavailable for a new customer. A high row-split score can therefore answer the wrong question.

An overlapping-customer split is not automatically wrong for every task. If you predict future activity for existing customers, a time-respecting split may be appropriate. In that case, every feature must still be available at the prediction timestamp; future labels, aggregates and lookups must not enter training.

Code / output
Target: unseen customers
customer IDs -> disjoint groups -> fit on training group only
                               -> evaluate on held-out group

See the mismatch without training a complex model

Python 3, standard library only. A lookup memorizes customer labels. The example is intentionally constructed, so its percentages illustrate the failure mode rather than estimate real model quality.

Code / output
rows = [
    {"customer": customer, "repeat": repeat, "label": customer % 2}
    for customer in range(8)
    for repeat in range(2)
]

def fit(train):
    return {row["customer"]: row["label"] for row in train}

def accuracy(model, test):
    predictions = [
        model.get(row["customer"], 0) == row["label"]
        for row in test
    ]
    return sum(predictions) / len(predictions)

row_train = [row for row in rows if row["repeat"] == 0]
row_test = [row for row in rows if row["repeat"] == 1]
group_train = [row for row in rows if row["customer"] < 4]
group_test = [row for row in rows if row["customer"] >= 4]

overlap = {r["customer"] for r in row_train} & {
    r["customer"] for r in row_test
}
assert len(overlap) == 8
assert not (
    {r["customer"] for r in group_train}
    & {r["customer"] for r in group_test}
)
assert accuracy(fit(row_train), row_test) == 1.0
assert accuracy(fit(group_train), group_test) == 0.5
print("row split:", accuracy(fit(row_train), row_test))
print("group split:", accuracy(fit(group_train), group_test))

Expected output:

Code / output
row split: 1.0
group split: 0.5

The row example deterministically alternates repeated records rather than drawing a random split, to make the overlap easy to inspect. Random row splitting can create the same overlap, but its exact amount depends on sampling.

Fix more than the split

  1. Decide whether the holdout unit is a customer, patient, document, organization, session or time period. Record that deployment assumption before tuning.
  2. Make the groups disjoint for unseen-entity evaluation. For repeated experiments, use group-aware cross-validation and inspect class balance across folds.
  3. Learn imputers, scalers, feature selection and encoders on each training fold only. A pipeline prevents many preprocessing mistakes, but it cannot repair leaked features or the wrong split strategy.
  4. Keep the final test set out of prompt, feature and model selection. Tuning repeatedly on it makes it part of development.
  5. Inspect duplicates and near-duplicates across partitions. Changing a record identifier does not make copied content independent.

For retrieval-augmented generation (RAG), the analogous question is whether near-identical questions, document versions or answer-bearing artifacts cross your chosen evaluation boundary. This does not mean all retrieval documents must be excluded: production RAG is supposed to retrieve useful knowledge. Define what information should be available and what must remain unseen for the task you are measuring.

Evidence to include in a follow-up question

Share group counts per split, overlap checks, timestamps, the feature-availability rule and how preprocessing was fitted. An unexplained score alone cannot establish leakage or prove that it has been eliminated.

Source checked 2026-10-07: scikit-learn: Common pitfalls and data leakage. Its key guidance is to split before fitting preprocessing and never learn transformations from the test set. The synthetic entity-overlap demonstration and RAG analogy are our AI-assisted explanation, not benchmark results.

0 score

Sign in to vote or contribute. Scores count actual member votes, not editorial endorsements.

Verify your sign-in to contribute an answer

Reading is free. Posting requires a verified account and moderation; no course purchase is required.