QuiddityML

Blog

ML concepts, explained so they stick

ML concepts explained in plain language, one idea at a time, with the math and the code side by side.

Start here

Browse by track

Learning and careers How to learn ML, what to learn first, and how to get hired. 7 posts ML Foundations How models learn, neural networks, optimization, generalization, and evaluation. 29 posts
Math for ML Linear algebra, calculus, probability, and information theory, as ML uses them. Coming soon
NLP From text and tokenization to embeddings, transformers, and LLMs. Coming soon
Computer Vision Images as data, CNNs, vision transformers, detection, segmentation, and more. Coming soon

All posts

Optimization Batch vs stochastic vs mini-batch gradient descent, and how to choose a batch size Batch, stochastic, and mini-batch gradient descent differ only in how many training examples each update looks at. This post explains what that changes, why mini-batch is the usual choice, and how to pick a batch size and adjust the learning rate with it. 30 September 2026 · 7 min read Optimization Learning rate schedules explained: warmup, cosine annealing, step decay, and ReduceLROnPlateau A learning rate schedule changes the step size of training as the run goes on, usually large at the start and small at the end. This post explains why a fixed learning rate falls short, how linear warmup, cosine annealing, step decay, and ReduceLROnPlateau work, and how to set each one up in PyTorch. 30 September 2026 · 12 min read Optimization Train, validation, and test sets: what each is for and how to split A dataset is usually split into three parts: a training set the model learns from, a validation set for making choices, and a test set for the final score. This post explains what each split is for, why the test set is kept out of model choices, and how to split data in PyTorch. 30 September 2026 · 7 min read Neural networks Dying ReLU: what dead neurons are and how to fix them A dead ReLU neuron outputs zero for every input and stops learning, and a network can lose a large share of its neurons this way during training. This post explains why ReLU neurons die, how to measure the dead fraction in PyTorch, and how to prevent it with initialization, the learning rate, and Leaky ReLU. 29 September 2026 · 8 min read Neural networks Vanishing and exploding gradients, and how gradient clipping helps Vanishing and exploding gradients happen when the training signal shrinks to almost nothing or grows out of control as it passes back through a deep network. This post shows why both happen, how to measure them in PyTorch, and how ReLU, good initialization, and gradient clipping help. 29 September 2026 · 7 min read Neural networks Weight initialization explained: Xavier vs He (Kaiming) vs orthogonal Weight initialization is how a neural network's weights are set before training starts, and a bad choice can stop a deep network from learning at all. This post shows what goes wrong with zero, too large, and too small weights, and when to use Xavier, He (Kaiming), or orthogonal initialization in PyTorch. 29 September 2026 · 8 min read Neural networks Parameters vs hyperparameters in machine learning Parameters are the numbers a model learns from data, and hyperparameters are the settings you choose before training starts. This post explains the difference with a small PyTorch network, shows where each one lives in code, and how to pick hyperparameters with a validation set. 28 September 2026 · 6 min read Neural networks The universal approximation theorem explained (and what it does not promise) The universal approximation theorem says a neural network with one wide enough hidden layer can match any continuous curve as closely as you like. This post explains what that means, why it works, and the four things the theorem does not promise about training and real data. 28 September 2026 · 6 min read Neural networks What is a neural network? The multilayer perceptron explained A neural network passes its input through layers of simple units, each layer reshaping the numbers for the next one. This post explains the multilayer perceptron layer by layer, why stacking layers helps, and how to count its parameters by hand and in PyTorch. 28 September 2026 · 8 min read How a model learns Logistic regression explained: sigmoid, binary cross-entropy, and why it is a one-layer neural network Logistic regression predicts the probability that an input belongs to one of two classes. This post explains the sigmoid function, the binary cross-entropy loss it trains with, why the two are paired, and how to train it in PyTorch. 27 September 2026 · 9 min read How a model learns Multi-class vs multi-label classification: softmax vs sigmoid outputs Multi-class classification picks exactly one label per input, and multi-label classification can pick several or none. This post explains the difference, why the first uses softmax and the second uses one sigmoid per label, and how to write both in PyTorch. 27 September 2026 · 5 min read How a model learns Softmax explained: probabilities, model confidence, and the stable softmax trick Softmax turns a classifier's raw scores into probabilities that add up to 1. This post explains the formula with a worked example, how to read its output as the model's confidence and when not to trust it, and why softmax code subtracts the largest score first. 27 September 2026 · 7 min read How a model learns Entropy, cross-entropy, and KL divergence explained Cross-entropy is the usual loss for training a classifier. This post explains where it comes from, starting with entropy and ending with KL divergence, works each one out on the same small example, and shows how to compute them in PyTorch. 26 September 2026 · 8 min read How a model learns Maximum likelihood estimation explained: where MSE and cross-entropy come from Maximum likelihood estimation picks the model parameters that make the observed data most probable. This post explains it with a coin-flip example, then shows how the same idea produces mean squared error for regression and cross-entropy for classification. 26 September 2026 · 8 min read How a model learns The PyTorch training loop explained, with Dataset and DataLoader The training loop is the five lines of PyTorch code that make a model learn from data. This post explains what each line does, why their order can't be swapped, and how a Dataset and DataLoader feed the loop one batch at a time. 26 September 2026 · 8 min read How a model learns Linear regression from scratch in PyTorch Linear regression predicts a number, like a house price, as a weighted sum of the inputs plus a bias. This post builds it in PyTorch two ways: with gradients computed by hand, and with nn.Linear and an optimizer. 25 September 2026 · 6 min read How to learn ML Learning ML with Python: the libraries you need (NumPy, pandas, scikit-learn, PyTorch) and what each is for Four Python libraries come up when you start machine learning. This post explains what NumPy, pandas, scikit-learn and PyTorch each do, which one to learn first, and when you need the other three. 25 September 2026 · 6 min read How a model learns What is a perceptron? From the original perceptron to the modern neuron A perceptron is the simplest artificial neuron: it weighs its inputs, adds them up, and outputs 0 or 1. This post explains how it works, how it learns, why modern neurons output a real number instead, and why one neuron can never solve XOR. 25 September 2026 · 7 min read Interviews and careers Machine learning interview questions for beginners, with answers About 30 machine learning questions that come up in junior ML, data science, and ML engineering interviews, each with a short answer and the follow-up an interviewer usually asks next, grouped from data and evaluation through training, metrics, neural networks, and architectures. 24 September 2026 · 19 min read How a model learns What is machine learning? The 5 types of learning explained Machine learning is a way to build software that learns its rules from examples instead of having them written by hand. This post explains how that works and why it can work on data the model has never seen, then walks through the five types of learning (supervised, unsupervised, semi-supervised, self-supervised, and reinforcement learning) with an example of each and how to pick between them. 24 September 2026 · 8 min read Interviews and careers How to choose a machine learning project A good machine learning project usually starts from a problem you or someone close to you actually has. This post covers where to find one, the questions to ask before you start, and what a project needs to be worth showing: a baseline, more than one dataset, more than one model, and more than one way to evaluate it. 23 September 2026 · 7 min read How a model learns What is a loss function? MSE, cross-entropy, and when to use which A loss function turns a model's predictions into one number that says how wrong they are, and training is the process of pushing that number down. This post explains why models need one, which loss to use for regression, binary, multi-class, and multi-label problems (and for a few other tasks), and how to use the loss on a proper validation split to evaluate a model. 23 September 2026 · 13 min read Neural networks Activation functions roadmap: step, sigmoid, tanh, ReLU, Leaky ReLU, GELU An activation function is the small non-linear function each neuron applies to its output. This post goes through the common ones in the order they were invented, what each fixed, what each broke, and how to write every one in PyTorch. 22 September 2026 · 6 min read Generalization and regularization Bias-variance tradeoff explained with pictures A model can be wrong on new data in two different ways, and the fixes for each pull in opposite directions. This post shows both with plots, how to tell them apart from your training and validation loss, and what to change for each. 22 September 2026 · 7 min read Interviews and careers What ML projects actually get you hired in 2026? The machine learning projects that help in hiring are the ones you cared enough about to finish, measure, and defend. This post covers how to pick one, what a reviewer checks, five projects worth building, how to present it, and how to answer when an interviewer asks why you didn't do it another way. 21 September 2026 · 7 min read Generalization and regularization Regularization explained: L1, L2, dropout, and early stopping Regularization is a set of techniques that stop a model from memorizing its training data so it does better on new data. This post explains four common ones (L2 weight decay, L1, dropout, and early stopping), how each one works, and how to set them in PyTorch. 21 September 2026 · 6 min read How a model learns What is backpropagation? Explained with a tiny network Backpropagation is how a neural network works out which way to change each of its weights to reduce its error. This post computes it by hand on a network with two weights, shows why it runs backward, and checks the numbers against PyTorch. 21 September 2026 · 5 min read How to learn ML How much math do you actually need for machine learning? You need enough math to read a loss function, an update rule, and a shape error, and that is a list you can finish in a few weeks. This post names every piece, says which ones you need before your first model and which can wait, and estimates the hours. 17 September 2026 · 5 min read Optimization Machine learning optimizers explained: SGD, momentum, RMSProp, Adam, AdamW An optimizer is the rule that turns gradients into parameter updates. This post explains the five optimizers most models train with, in the order each one was invented to fix the last, and which one to pick. 17 September 2026 · 7 min read Generalization and regularization Overfitting vs underfitting: how to tell which one you have, and how to fix it Overfitting is when a model memorizes its training data and fails on new data. Underfitting is when it cannot even fit the training data. This post shows how to read which one you have from the loss curves and what to change for each. 17 September 2026 · 6 min read How to learn ML How to learn machine learning effectively (and actually remember it) Most people finish an ML course and cannot use it a month later. This post covers what to learn in which order, how to practice so the concepts stay, and a weekly schedule you can keep with 30 to 45 minutes a day. 16 September 2026 · 7 min read Evaluation My loss is not decreasing: a checklist A training loss that stays flat usually has one of a short list of causes. This checklist goes through them in a practical order: the data, the training loop, the learning rate, the loss function, initialization, and a one-batch test that separates a bug from a hard problem. 16 September 2026 · 7 min read Optimization What is a learning rate, and how do you pick one? The learning rate is the number that most often decides whether a model trains well. This post explains what it does inside the update rule, how to read the three shapes of a loss curve, why good values differ between SGD and Adam, and a five-line test that finds a usable value. 16 September 2026 · 6 min read How a model learns What is gradient descent? Explained step by step Gradient descent is the procedure that trains almost every machine learning model. This post explains what it does, why the update has a minus sign, how the learning rate changes everything, and what the three lines of PyTorch that implement it are doing. 16 September 2026 · 6 min read Optimization What is the Adam optimizer? Explained step by step Adam is the optimizer most deep learning models train with. This post explains what it does, why it works better than plain gradient descent, and when to use AdamW instead. 15 September 2026 · 5 min read How to learn ML How to learn machine learning in 2026: a complete roadmap for self-taught learners The order to learn machine learning on your own: Python, the math that ML uses, PyTorch, then training and evaluating deep networks, with every concept named so you know exactly what is ahead of you. 15 September 2026 · 14 min read