30 September 2026 · 7 min read
Train, validation, and test sets: what each is for and how to split
A dataset is usually split into three parts: a training set the model learns from, a validation set for making choices, and a test set for the final score. This post explains what each split is for, why the test set is kept out of model choices, and how to split data in PyTorch.
Machine learning data is usually split into three separate parts. The training set is what the model learns from, the validation set is what you check while making choices like the learning rate or the model size, and the test set is held back until the end to give one final score on data that played no part in building the model. The split exists because a score measured on the training data says how well the model fits examples it has already seen, while the question people care about is how it does on new ones.
Why the training score is not enough
Training a model means adjusting its weights so that a loss, one number measuring how wrong the predictions are, goes down on the training examples. That loss is measured on the same examples the weights were tuned to.
A model with enough weights can push that loss close to zero by memorizing the training examples: it stores which input goes with which label instead of learning a rule that carries over to inputs it has not seen. This is called overfitting. A memorizing model and a model that learned a real rule can have the same training loss, and the training loss alone has no way to tell them apart.
Separating them takes a score measured on examples the model did not train on. Those held-out examples play the part of new data. A model that learned a rule scores well on them, and a model that memorized scores near chance.
What each split is for
- Training set: the examples used to compute the loss and its gradient at each training step. The weights are fit to this data.
- Validation set: held-out examples used during development, for any decision a person or a script makes. Which learning rate, how many layers, how wide each layer is, when to stop training, which of two architectures to keep. Hyperparameters, the settings chosen before training rather than learned, are picked by comparing validation scores.
- Test set: held-out examples used once, after the model is final, to estimate how the finished model does on new data.
A common starting split is 70% training, 15% validation and 15% test.

Why the test set stays out of model choices
The validation and test sets are both held-out data, so it can look like one of them is redundant. The difference is what happens to a score once it has been used to make a choice.
Say you train 40 models with different settings and keep the one with the highest validation accuracy. Each validation score is the model's real skill plus some luck from which 150 examples happened to land in the validation set. Picking the highest of 40 scores tends to pick one with a lot of luck in it, so the winner's validation score usually comes out higher than what the model will get on new data. The more settings you try, the larger that gap tends to get.
The test set avoids this by taking no part in the choice. Its score has no selection luck built in, so it stays a fair estimate. If you look at the test score after each run and keep the model that does best there, the test set has become a second validation set, and its score is now optimistic in the same way.
The experiment below makes the gap visible. The labels are random coin flips, so no model can do better than 50% on new data on average. It trains 40 small networks with different widths, learning rates and random seeds, keeps the one with the best validation accuracy, and then checks it on the test set.
1import torch
2import torch.nn as nn
3
4torch.manual_seed(0)
5X = torch.randn(1000, 16)
6y = torch.randint(0, 2, (1000,)).float() # random labels: nothing to learn
7
8perm = torch.randperm(1000)
9tr, va, te = perm[:700], perm[700:850], perm[850:]
10
11def accuracy(model, idx):
12 with torch.no_grad():
13 return ((model(X[idx]).squeeze(1) > 0).float() == y[idx]).float().mean().item()
14
15def train(width, lr, seed):
16 torch.manual_seed(seed)
17 model = nn.Sequential(nn.Linear(16, width), nn.ReLU(), nn.Linear(width, 1))
18 opt = torch.optim.Adam(model.parameters(), lr=lr)
19 for _ in range(300):
20 loss = nn.functional.binary_cross_entropy_with_logits(model(X[tr]).squeeze(1), y[tr])
21 opt.zero_grad()
22 loss.backward()
23 opt.step()
24 return model
25
26trials = []
27for width in [8, 32, 128, 512]:
28 for lr in [0.001, 0.01]:
29 for seed in range(5):
30 m = train(width, lr, seed)
31 trials.append((accuracy(m, va), width, lr, seed, m))
32
33val_acc, width, lr, seed, best = max(trials, key=lambda t: t[0])
34print(f"trials: {len(trials)}, mean val accuracy: {sum(t[0] for t in trials) / len(trials):.3f}")
35print(f"winner: width={width}, lr={lr}, seed={seed}")
36print(f"train accuracy: {accuracy(best, tr):.3f}")
37print(f"val accuracy: {val_acc:.3f}")
38print(f"test accuracy: {accuracy(best, te):.3f}")1trials: 40, mean val accuracy: 0.518
2winner: width=32, lr=0.01, seed=4
3train accuracy: 1.000
4val accuracy: 0.587
5test accuracy: 0.513The winning model gets 100% on the training set because it memorized 700 random labels. Its validation accuracy of 58.7% looks like it found something, but that number was picked as the highest of 40, while the average validation accuracy across the 40 models was 51.8%. The test set, which played no part in the choice, reports 51.3%, close to the 50% that random labels allow.
How to split a dataset in PyTorch
For data stored in tensors, shuffle the row indices once with torch.randperm and slice them. For a PyTorch Dataset, torch.utils.data.random_split does the shuffling and slicing for you. In both cases a fixed seed keeps the split the same between runs, so a change in the validation score comes from a change in the model rather than from a new set of validation examples.
1import torch
2from torch.utils.data import TensorDataset, random_split
3
4torch.manual_seed(0)
5X = torch.randn(1000, 16)
6y = torch.randint(0, 2, (1000,))
7
8# Option 1: shuffle the row indices once, then slice them
9perm = torch.randperm(len(X))
10n_train, n_val = int(0.7 * len(X)), int(0.15 * len(X))
11train_idx = perm[:n_train]
12val_idx = perm[n_train:n_train + n_val]
13test_idx = perm[n_train + n_val:]
14print(len(train_idx), len(val_idx), len(test_idx))
15
16# No row appears in two splits
17overlap = set(train_idx.tolist()) & (set(val_idx.tolist()) | set(test_idx.tolist()))
18print("overlap:", len(overlap))
19
20# Option 2: random_split on a Dataset, with a seeded generator
21dataset = TensorDataset(X, y)
22train_set, val_set, test_set = random_split(
23 dataset, [0.7, 0.15, 0.15],
24 generator=torch.Generator().manual_seed(42),
25)
26print(len(train_set), len(val_set), len(test_set))1700 150 150
2overlap: 0
3700 150 150The test size is computed as whatever is left over (perm[n_train + n_val:]) instead of as int(0.15 * len(X)), so rounding cannot drop a row. When the sizes print wrong, check that the fractions add up to 1 first.
Preprocessing with training statistics
Normalizing features means subtracting a mean and dividing by a standard deviation. Those two numbers are computed from the training split and then applied unchanged to the validation and test splits. Computing them on the full dataset lets information from the held-out rows into training.
1import torch
2
3torch.manual_seed(0)
4X = 5 + 2 * torch.randn(1000, 16)
5perm = torch.randperm(len(X))
6X_train, X_val, X_test = X[perm[:700]], X[perm[700:850]], X[perm[850:]]
7
8# Statistics come from the training split only
9mean = X_train.mean(dim=0)
10std = X_train.std(dim=0)
11
12X_train = (X_train - mean) / std
13X_val = (X_val - mean) / std # same mean and std, not recomputed
14X_test = (X_test - mean) / std
15print(f"train mean {X_train.mean():.3f}, val mean {X_val.mean():.3f}, test mean {X_test.mean():.3f}")1train mean -0.000, val mean -0.028, test mean -0.017The validation and test means are close to zero without being exactly zero, which is expected: they are new rows run through a transform fit on other rows.
How big each split should be
70/15/15 is a starting point, and the right sizes depend mostly on how many examples you have. What the validation and test sets need is enough examples for their scores to be stable, and past that point extra held-out examples mostly take data away from training. With 1 million examples, a 1% test set is 10,000 examples, which is plenty for most accuracy estimates. With 1,000 examples, holding out 20 to 30% is common, and a 150-example validation set is already noisy enough that one lucky model can stand out by several points, the way the 58.7% winner did. When data is that scarce, cross-validation, which rotates which part of the data is held out and averages the scores, gets more out of it.
Common mistakes
- Tuning on the test set. Checking the test score after each run and keeping the best one turns it into a validation set. Look at it once, after the model is final.
- Splitting at random when rows are related. If one patient, user or document shows up in several rows, a random split can put some of their rows in training and others in the test set, and the test score rewards the model for recognizing them. Split by patient, user or document instead. For data ordered in time, train on the past and validate and test on later periods.
- Fitting preprocessing on the full dataset. Means, standard deviations and vocabularies come from the training split.
- Resplitting on each run. Without a fixed seed,
random_splitdraws a new split each time, and two runs are scored on different validation examples. - Checking the class balance too late. With a rare class, a random split can leave the validation set with only a handful of positives. Count the labels in each split before training.
Related concepts
Overfitting and underfitting are what the gap between training and validation scores measures. Parameters vs hyperparameters walks through a learning-rate sweep scored on a validation set with a single test check at the end. Regularization covers early stopping, which uses the validation loss to decide when training ends. Cross-validation is the alternative to a single validation split when the dataset is small.
QuiddityML teaches train, validation and test splits as a concept in Unit 3 of the ML Foundation track, where the exercises include writing a seeded split_dataset function with random_split, spotting the version with no generator, and tracing the split sizes and batch shapes of its DataLoaders (quiddityml.com).