0 score1 answer
Why can a random train/test split look strong but fail on new customers?
A dataset has several rows per customer. A row-level random split places records for the same customer in both training and evaluation. A model performs well in that test but poorly for customers it has never seen. Is the split measuring the intended deployment task, and what should change? This ed