QuiddityML

30 September 2026 · 12 min read

Learning rate schedules explained: warmup, cosine annealing, step decay, and ReduceLROnPlateau

A learning rate schedule changes the step size of training as the run goes on, usually large at the start and small at the end. This post explains why a fixed learning rate falls short, how linear warmup, cosine annealing, step decay, and ReduceLROnPlateau work, and how to set each one up in PyTorch.

A learning rate schedule is a rule that changes the learning rate while a model trains, instead of keeping it at one value for the whole run. The usual pattern is a large rate early, when the model is far from a good solution, and a small rate late, when it is close to one. Many training scripts use a schedule, and picking the wrong one, or wiring it up wrong, can leave a model at a higher loss than the same code with a better schedule.

What the learning rate does

Training updates a model's parameters $\theta$ over and over. On each step it computes the gradient $\nabla L(\theta)$, the direction in which the loss $L$ grows fastest, and moves the parameters a small distance the other way:

$$\theta \leftarrow \theta - \eta , \nabla L(\theta)$$

The learning rate $\eta$ sets how far each step goes. A gradient of 2.0 with $\eta = 0.1$ moves a parameter by 0.2. How to pick a starting value is covered in what a learning rate is and how to pick one. A schedule starts from that value and changes it as training goes on, so $\eta$ becomes $\eta_t$, the learning rate at step $t$.

Why a fixed learning rate is a compromise

Training has two phases that want different step sizes.

A single fixed value has to pick one of these. A large rate makes fast early progress and then bounces at the end. A small rate settles well at the end but spends most of the run creeping forward. Starting high and lowering the rate over time gets the fast start and the precise finish in the same run, which is what learning rate scheduling is.

Top: gradient descent steps on a loss contour, large steps early covering ground and small steps late refining the solution near the minimum. Bottom: a learning rate that starts high and decays low over iterations.

How a scheduler works in PyTorch

In PyTorch the learning rate lives in the optimizer, in optimizer.param_groups[0]["lr"]. A scheduler is an object that wraps the optimizer and rewrites that value each time scheduler.step() is called. Depending on the schedule, scheduler.step() is called once per batch or once per epoch (one full pass over the training data), and the schedule's lengths, such as T_max or step_size, count those calls.

Linear warmup

The first steps of training are the least stable. The weights are random, so the gradients they produce are large and point in inconsistent directions. Adaptive optimizers such as Adam add a second problem: Adam divides each parameter's step by the square root of $\hat{v}$, a running average of that parameter's squared gradients. At step 5, $\hat{v}$ has been estimated from five gradients, so it is noisy, and dividing by a noisy number gives erratic step sizes. A full-sized learning rate on top of that can push the weights into a bad region early in the run, and the run may not recover from it.

Linear warmup starts the learning rate near zero and raises it in a straight line to its peak value $\eta_{\max}$ over the first $T_{\text{warmup}}$ steps:

$$\eta_t = \eta_{\max} \cdot \frac{t}{T_{\text{warmup}}} \quad \text{for } t < T_{\text{warmup}}$$

At $t = 0$ the rate is 0, halfway through warmup it is $\eta_{\max}/2$, and at $t = T_{\text{warmup}}$ it reaches the peak. By then the gradients are less erratic and Adam's $\hat{v}$ has been averaged over hundreds of steps. Warmup lengths from a few hundred to a few thousand steps are common. Large models and large batch sizes tend to need it most, since one bad early update can undo a good initialization.

Linear warmup: the learning rate rises in a straight line from 0 to the peak at T_warmup, then the rest of the schedule continues downward.

Cosine annealing

After warmup, the learning rate needs to come down. Cosine annealing lowers it along half of a cosine curve, from the peak $\eta_{\max}$ to a floor $\eta_{\min}$ at the last step $T$:

$$\eta_t = \eta_{\min} + \frac{1}{2}(\eta_{\max} - \eta_{\min})\left(1 + \cos\left(\pi \cdot \frac{t - T_{\text{warmup}}}{T - T_{\text{warmup}}}\right)\right)$$

The fraction inside the cosine is how far through the decay phase the run is, going from 0 right after warmup to 1 at the end. Over that range $\cos$ goes from 1 to -1, so the factor $\frac{1}{2}(1 + \cos)$ goes from 1 to 0 and the rate moves from $\eta_{\max}$ down to $\eta_{\min}$. The curve is flat at both ends and steepest in the middle: the rate stays near the peak for a while after warmup, drops fastest halfway through, and eases into the floor at the end. There are no sudden jumps, and the final steps run at a rate close to $\eta_{\min}$, which is where the precise settling happens.

Warmup followed by cosine annealing: the rate climbs linearly from eta_min to eta_max at T_warmup, then follows a cosine curve down to eta_min at step T, with the formula and three properties: smooth, slow-fast-slow, and near zero at the end.

The two formulas together fit in one short function:

1import math
2 
3def get_lr(step, peak_lr, warmup_steps, total_steps, min_lr=1e-6):
4    """Linear warmup to peak_lr, then cosine decay to min_lr."""
5    if step < warmup_steps:
6        return peak_lr * step / warmup_steps
7    progress = (step - warmup_steps) / (total_steps - warmup_steps)
8    return min_lr + 0.5 * (peak_lr - min_lr) * (1 + math.cos(math.pi * progress))
9 
10for step in [0, 50, 100, 325, 550, 775, 1000]:
11    print(f"step {step:4d}  lr {get_lr(step, 1e-3, 100, 1000):.2e}")
1step    0  lr 0.00e+00
2step   50  lr 5.00e-04
3step  100  lr 1.00e-03
4step  325  lr 8.54e-04
5step  550  lr 5.01e-04
6step  775  lr 1.47e-04
7step 1000  lr 1.00e-06

With a peak of 1e-3, 100 warmup steps, and 1,000 steps in total, the rate is half the peak at step 50, reaches the peak at step 100, is back to half the peak at step 550 (the middle of the decay phase), and ends at the floor of 1e-6.

Warmup and cosine in PyTorch

PyTorch builds the same schedule from three pieces. LinearLR multiplies the base learning rate by a factor that moves from start_factor to end_factor over total_iters steps, which gives the warmup. CosineAnnealingLR does the cosine decay over T_max steps down to eta_min. SequentialLR runs the first scheduler until the step count reaches the milestones value, then hands over to the second.

1import torch
2from torch.optim.lr_scheduler import LinearLR, CosineAnnealingLR, SequentialLR
3 
4model = torch.nn.Linear(10, 1)
5optimizer = torch.optim.AdamW(model.parameters(), lr=1e-3)
6 
7warmup_steps, total_steps = 100, 1000
8warmup = LinearLR(optimizer, start_factor=1e-6, end_factor=1.0, total_iters=warmup_steps)
9cosine = CosineAnnealingLR(optimizer, T_max=total_steps - warmup_steps, eta_min=1e-6)
10scheduler = SequentialLR(optimizer, schedulers=[warmup, cosine], milestones=[warmup_steps])
11 
12X, y = torch.randn(256, 10), torch.randn(256, 1)
13for step in range(total_steps):
14    lr = optimizer.param_groups[0]["lr"]
15    if step in (0, 50, 100, 325, 550, 775, 999):
16        print(f"step {step:4d}  lr {lr:.2e}")
17    loss = torch.nn.functional.mse_loss(model(X), y)
18    optimizer.zero_grad()
19    loss.backward()
20    optimizer.step()
21    scheduler.step()
1step    0  lr 1.00e-09
2step   50  lr 5.00e-04
3step  100  lr 1.00e-03
4step  325  lr 8.54e-04
5step  550  lr 5.00e-04
6step  775  lr 1.47e-04
7step  999  lr 1.00e-06

The printed rates agree with get_lr to two significant figures, and the last line is step 999 because the loop runs steps 0 to 999. The one visible difference is step 0, where LinearLR starts at start_factor times the base rate, 1e-9, instead of exactly 0. T_max is total_steps - warmup_steps because the cosine part only runs after warmup ends.

scheduler.step() goes after optimizer.step(). If the order is reversed, PyTorch prints a warning and skips the first value of the schedule, so each optimizer step runs one step ahead of the schedule and the first step uses the wrong rate.

Warmup then cosine in code: SequentialLR hands over from LinearLR (start_factor 1e-6, total_iters = warmup_steps) to CosineAnnealingLR (T_max = total - warmup, eta_min = 1e-6) at the warmup_steps milestone. The correct order is loss.backward(), optimizer.step(), scheduler.step().

Step decay

Step decay is an older and simpler schedule: keep the rate fixed, and multiply it by a factor $\gamma$ (gamma, a number below 1) every $T_{\text{step}}$ epochs. Starting from $\eta_0$:

$$\eta_t = \eta_0 \cdot \gamma^{\lfloor t / T_{\text{step}} \rfloor}$$

$\lfloor t / T_{\text{step}} \rfloor$ rounds down, so it counts how many full blocks of $T_{\text{step}}$ epochs have passed, and each block multiplies the rate by $\gamma$ once more. A common setting for image classifiers trained with SGD was $\eta_0 = 0.1$ and $\gamma = 0.1$ every 30 epochs. StepLR implements it and is stepped once per epoch:

1import torch
2 
3model = torch.nn.Linear(10, 1)
4optimizer = torch.optim.SGD(model.parameters(), lr=0.1, momentum=0.9)
5scheduler = torch.optim.lr_scheduler.StepLR(optimizer, step_size=30, gamma=0.1)
6 
7for epoch in range(100):
8    if epoch in (0, 29, 30, 59, 60, 90):
9        print(f"epoch {epoch:3d}  lr {optimizer.param_groups[0]['lr']:.0e}")
10    # ... one epoch of training batches, each ending in optimizer.step() ...
11    optimizer.step()
12    scheduler.step()
1epoch   0  lr 1e-01
2epoch  29  lr 1e-01
3epoch  30  lr 1e-02
4epoch  59  lr 1e-02
5epoch  60  lr 1e-03
6epoch  90  lr 1e-04

The rate holds at 0.1 through epoch 29, then drops by 10 times at epochs 30, 60, and 90. The loss usually falls sharply right after each drop and then flattens until the next one. The drawback is that the drop epochs are extra settings to choose by hand, and they have to be chosen again when the length of the run changes. Cosine annealing needs only the total length, which is one reason it has replaced step decay in many newer training setups. When exact step points are needed, MultiStepLR takes a list of epochs such as milestones=[30, 60, 90].

ReduceLROnPlateau

Cosine annealing and step decay both need to know in advance how long training will run. ReduceLROnPlateau does not. It watches a metric measured on a validation set, a slice of data held back from training so the metric says how the model does on examples it has not fit, and cuts the learning rate when that metric stops improving. Its main settings:

ReduceLROnPlateau with mode min, factor 0.1, patience 5 and min_lr 1e-6: the learning rate, on a log scale, steps down each time the validation loss flattens for the patience window, while the validation loss keeps falling after each cut.

Unlike the other schedulers, its step() takes the metric as an argument:

1import torch
2 
3model = torch.nn.Linear(10, 1)
4optimizer = torch.optim.SGD(model.parameters(), lr=0.1)
5scheduler = torch.optim.lr_scheduler.ReduceLROnPlateau(
6    optimizer, mode="min", factor=0.1, patience=2, min_lr=1e-6
7)
8 
9val_losses = [0.90, 0.70, 0.60, 0.61, 0.60, 0.62, 0.63, 0.50, 0.49, 0.49, 0.50, 0.50]
10for epoch, val_loss in enumerate(val_losses):
11    optimizer.step()  # stands in for one epoch of training
12    scheduler.step(val_loss)
13    print(f"epoch {epoch:2d}  val_loss {val_loss:.2f}  lr {optimizer.param_groups[0]['lr']:.0e}")
1epoch  0  val_loss 0.90  lr 1e-01
2epoch  1  val_loss 0.70  lr 1e-01
3epoch  2  val_loss 0.60  lr 1e-01
4epoch  3  val_loss 0.61  lr 1e-01
5epoch  4  val_loss 0.60  lr 1e-01
6epoch  5  val_loss 0.62  lr 1e-02
7epoch  6  val_loss 0.63  lr 1e-02
8epoch  7  val_loss 0.50  lr 1e-02
9epoch  8  val_loss 0.49  lr 1e-02
10epoch  9  val_loss 0.49  lr 1e-02
11epoch 10  val_loss 0.50  lr 1e-02
12epoch 11  val_loss 0.50  lr 1e-03

The lowest loss so far is 0.60 at epoch 2. Epochs 3, 4, and 5 do not beat it (matching 0.60 at epoch 4 does not count as an improvement), so after epoch 5 the count is 3, which is above patience=2, and the rate drops to 0.01. The count resets after the cut. Epochs 7 and 8 set new bests, then epochs 9, 10, and 11 fail to beat 0.49, and the rate drops again to 0.001.

ReduceLROnPlateau has no warmup built in, and it needs a validation pass at regular intervals, which costs extra compute on large runs. It is a common choice for smaller models and for runs whose length is not known ahead of time.

Does a schedule change the result?

The experiment below trains a linear model on 2,000 noisy examples with plain SGD, a batch of 16 random rows per step, and a starting rate of 0.05 for 2,000 steps. It compares a constant rate, step decay (divide by 10 every 600 steps), and cosine annealing, and measures the final loss on the full dataset. The last line is the lowest loss a linear model can reach on this data, solved directly with least squares.

1import torch
2 
3torch.manual_seed(0)
4X = torch.randn(2000, 20)
5y = X @ torch.randn(20, 1) + 0.5 * torch.randn(2000, 1)
6 
7def train(schedule, steps=2000, lr=0.05):
8    torch.manual_seed(1)
9    model = torch.nn.Linear(20, 1)
10    optimizer = torch.optim.SGD(model.parameters(), lr=lr)
11    scheduler = None
12    if schedule == "step":
13        scheduler = torch.optim.lr_scheduler.StepLR(optimizer, step_size=600, gamma=0.1)
14    elif schedule == "cosine":
15        scheduler = torch.optim.lr_scheduler.CosineAnnealingLR(optimizer, T_max=steps)
16    for _ in range(steps):
17        idx = torch.randint(0, 2000, (16,))  # a mini-batch of 16 random rows
18        loss = torch.nn.functional.mse_loss(model(X[idx]), y[idx])
19        optimizer.zero_grad()
20        loss.backward()
21        optimizer.step()
22        if scheduler is not None:
23            scheduler.step()
24    with torch.no_grad():
25        return torch.nn.functional.mse_loss(model(X), y).item()
26 
27for schedule in ["constant", "step", "cosine"]:
28    print(f"{schedule:8s}  final loss {train(schedule):.4f}")
29 
30# The lowest loss a linear model can reach on this data, solved directly
31Xb = torch.cat([X, torch.ones(2000, 1)], dim=1)
32w = torch.linalg.lstsq(Xb, y).solution
33print(f"least squares {torch.nn.functional.mse_loss(Xb @ w, y).item():.4f}")
1constant  final loss 0.2567
2step      final loss 0.2346
3cosine    final loss 0.2348
4least squares 0.2343

With a constant rate the model ends 0.0224 above the least-squares loss, because each noisy mini-batch step at 0.05 keeps moving the weights around the minimum. Step decay and cosine annealing both end within 0.0005 of it, since their last few hundred steps run at a rate small enough to settle. On this problem the two decaying schedules finish in the same place. The practical difference between them is that step decay needed a drop point (600 steps) chosen by hand, and cosine annealing needed only the total number of steps.

Which schedule to use

Schedule Needs total length Reacts to validation Common use
Warmup + cosine yes no transformers, most large runs
Step decay yes, to place drops no older SGD image recipes
ReduceLROnPlateau no yes small models, open-ended runs
Constant no no quick experiments, debugging

Common mistakes

The learning rate post covers picking the starting value that a schedule then shapes. Adam and the other optimizers decide the direction and per-parameter size of each step, and a schedule scales all of those steps together, which is why warmup and Adam are usually used as a pair. Early stopping is the close relative of ReduceLROnPlateau: both watch the validation metric, and one lowers the rate when it stalls while the other ends training. The training loop is where scheduler.step() sits, right after optimizer.step().

QuiddityML teaches each of these schedules as its own concept in the ML Foundation track, and the exercises include writing get_lr for warmup plus cosine from the formula, ordering the lines that set up SequentialLR, and tracing a ReduceLROnPlateau run epoch by epoch to find where the rate drops (quiddityml.com).