Multi-Layer Perceptrons & Backpropagation, Explained Once and For All
A single perceptron draws one straight line — and that's all it will ever do. The moment your data can't be sliced by a line, it fails, forever. Everything modern AI can do comes from two upgrades: stacking neurons into layers, and backpropagation — the algorithm that trains all those layers at once.
This is the second lesson in the neural networks series, following the perceptron explained. It mirrors the Multi-Layered Perceptron lesson from Microsoft's AI for Beginners curriculum — same concepts, same order, written for developers who'd rather read code than proofs.
Formalizing the learning problem
Machine learning in one sentence: given a training dataset X with labels Y, find a model f whose predictions minimize a loss function ℒ. The loss is how wrong you are — nothing more mystical than that.
Choosing the loss is a design decision:
| Problem | Typical loss | Formula |
|---|---|---|
| Regression | Absolute error | Σ|f(x⁽ⁱ⁾) − y⁽ⁱ⁾| |
| Regression | Squared error | Σ(f(x⁽ⁱ⁾) − y⁽ⁱ⁾)² |
| Classification | 0-1 loss | = accuracy, counted the other way |
| Classification | Logistic loss | smooth, trainable version of 0-1 |
One more piece: for classification we usually want probabilities, not raw scores. The softmax function σ converts arbitrary outputs into a probability distribution, making the model f(x) = σ(wx + b). The weights w and biases b together are the parameters θ, and training means minimizing ℒ by varying θ.
Gradient descent, now with layers
The optimizer is the same one from the perceptron lesson — initialize parameters randomly, then repeatedly step against the gradient:
w⁽ⁱ⁺¹⁾ = w⁽ⁱ⁾ − η ∂ℒ/∂w
b⁽ⁱ⁺¹⁾ = b⁽ⁱ⁾ − η ∂ℒ/∂b
η is still the learning rate. One practical upgrade: computing the gradient over the whole dataset is expensive, so real training uses minibatches — small random subsets — recalculated each step. That's stochastic gradient descent (SGD), and it's what every training loop you've ever run actually does.
The multi-layered perceptron
A two-layer network chains two linear maps with a non-linear activation α between them:
z₁ = w₁x + b₁
z₂ = w₂·α(z₁) + b₂
f = σ(z₂)
That α is the whole trick. Stack linear layers without a non-linearity between them and they collapse into one linear layer — you'd have built an expensive perceptron. With α between them, the network can bend its decision boundary around any shape of data. Parameters are now θ = ⟨w₁, b₁, w₂, b₂⟩, and gradient descent still applies — the only question is how to compute the gradients through the layers.
Backpropagation: the chain rule with a filing system
Here's the derivative of the loss with respect to an early-layer weight, via the chain rule:
∂ℒ/∂w₁ = (∂ℒ/∂σ)(∂σ/∂z₂)(∂z₂/∂α)(∂α/∂z₁)(∂z₁/∂w₁)
Look closely: the left-most factors are the same for every layer's gradient. So instead of recomputing them per weight, compute the derivatives once, starting from the loss and flowing backwards through the network, reusing each partial result. That reuse is the entire algorithm — backpropagation, or "backprop".
The computational graph makes it concrete:
x → [w₁x+b₁] → z₁ → [α(z₁)] → a₁ → [w₂a₁+b₂] → z₂ → [σ(z₂)] → f → [ℒ(f,y)] → loss
←──────────────── backprop: gradients flow this way ────────────────
Forward pass computes the loss. Backward pass multiplies local derivatives layer by layer back to the inputs. Every modern framework — PyTorch, TensorFlow, JAX — is at its core this bookkeeping, automated.
Training, end to end, in pseudocode
for epoch in range(num_epochs):
for batch in random_minibatches(X, Y, batch_size):
# forward pass
z1 = batch @ w1 + b1
a1 = relu(z1) # non-linear activation
z2 = a1 @ w2 + b2
f = softmax(z2)
loss = cross_entropy(f, batch_labels)
# backward pass (backprop)
grads = backprop(loss) # ∂ℒ/∂w₂, ∂ℒ/∂b₂, ∂ℒ/∂w₁, ∂ℒ/∂b₁
# SGD update
w2 -= eta * grads.w2; b2 -= eta * grads.b2
w1 -= eta * grads.w1; b1 -= eta * grads.b1
Four verbs, repeated: forward, backward, update, shuffle. That's all training is, at any scale — a 10-neuron toy and a frontier LLM differ in parameter count, not in kind.
Build it yourself once
The lesson's notebook has you implement this as a modular mini-framework — Linear layer, activation functions, loss functions, an optimizer class — then train it on a 2-D classification task where a single perceptron provably fails. The lab then scales the same framework to MNIST handwritten digits: ~98% accuracy from a network you wrote yourself.
That exercise is the fastest de-mystifier in deep learning. After it, "the model learned" stops being a metaphor: you have watched weights you initialized at zero organize themselves into digit recognition.
The series continues with moving from your own framework to real ones (PyTorch/TensorFlow) in the next lesson of Microsoft's AI for Beginners. For the video-first path through these same concepts, see the DevKingOv courses.
FAQ
Why do you need a non-linear activation between layers?
Without one, stacking linear layers is pointless: the composition of linear functions is itself linear, so a 100-layer linear network has exactly the expressive power of a single perceptron. The non-linearity is what lets networks approximate curved, entangled decision boundaries.
Is backpropagation how the brain learns?
Almost certainly not — it needs exact symmetric weights and a global error signal, which biological neurons don't have. Backprop is an efficient computational scheme for credit assignment, not a neuroscience claim. It won because it works, not because it's brain-shaped.
What is the vanishing gradient problem?
In deep networks, backprop multiplies many small derivatives on its way to early layers, so gradients can shrink toward zero — early layers stop learning. Smooth activations with better gradients (ReLU), careful initialization, normalization layers, and residual connections are the standard fixes.
How many layers do I actually need?
Start tiny. Many tabular problems are solved by 1–3 hidden layers; images typically want convolutional depth; language wants transformers. The lesson's MNIST lab reaches ~98% with one or two hidden layers — a reminder that architecture size should match problem size.
SGD, minibatch, batch — what's the difference?
Batch gradient descent uses all training data per step (stable but slow). Pure SGD uses one example per step (noisy but cheap). Minibatch SGD — a small random subset per step — is the industry default because it parallelizes on GPUs and the noise actually helps escape shallow local minima.
Where do I go after this lesson?
Implement the framework in the AI for Beginners notebook, then the same curriculum's lesson on real frameworks (PyTorch/TensorFlow), then CNNs for vision and embeddings/transformers for NLP. On this site, start with how to learn AI-assisted coding to place it in a learning path.
Prefer watching?
Every post here is a lesson in a free video course — follow along on YouTube and track your progress on the portal.
Keep reading
The Perceptron Explained — Where Neural Networks Began (1957 to Python)
The complete beginner's guide to the perceptron: the 1957 Mark-1 machine, threshold logic units, the step activation function, perceptron training, and gradient descent — with working Python code and no math background required.
Build a Document Q&A App — Your Second AI App (RAG in Plain JavaScript)
Let users upload documents and ask questions about them. This guide builds retrieval-augmented generation from first principles in plain JavaScript: chunking, embeddings, vector search, and cited answers — no ML background needed.
Build Your First AI App with JavaScript — A Beginner's Guide
The complete path to your first deployed AI-powered app: how AI APIs actually work, the safe server-route pattern, streaming responses, and which first project to pick. No machine learning required — just JavaScript.
Want help applying this? Book a 1-on-1 with a consultant.
Find a Consultant