DEEP LEARNING FRAMEWORK

PyTorch Tutorials

PyTorch is what most deep learning research is written in, and increasingly what production runs on too. These 32 tutorials cover tensors, layers, loss functions and the training loop you write yourself.

  • 32 tutorials
  • 5 topics
  • Explicit training loop
Five lines of loop, and nothing is hiddenRead the DataLoader tutorial

What PyTorch is for

PyTorch is Meta’s deep learning framework, and it has become the default in research. Its appeal is that it behaves like ordinary Python. Tensors work like NumPy arrays, models are normal classes, and the training loop is a for loop you write and can step through in a debugger.

That explicitness is the main difference from Keras. Nothing calls fit() for you: each batch, you run the model forward, compute the loss, call backward() to get gradients, and ask the optimizer to take a step. It is four more lines, and in exchange there is no hidden behaviour to fight when you need something unusual.

The one habit to build early is optimizer.zero_grad(). PyTorch accumulates gradients by design, which is useful for some techniques but means that forgetting to clear them silently ruins training. Beyond that, most day-to-day PyTorch work is shape wrangling — reshaping, adding a batch dimension, and stacking tensors — which is why that group is the largest below.

Install it and make a tensor

The right pip command depends on whether you want GPU support, so take it from the official selector rather than guessing — installing the CPU wheel and then wondering why training is slow is a common detour.

Once it imports, the first thing to learn is moving between NumPy and PyTorch. torch.from_numpy() and .numpy() convert in both directions and share memory, so changing one changes the other.

RuntimeError: Expected all tensors to be on the same device means part of your data is on the GPU and part is on the CPU. Move both with .to(device) — the model and every batch — using one device variable defined at the top of the script.

False just means CPU, which is fine to learn onNumPy conversion guide

If you are starting today

A sensible order to learn PyTorch in

Five steps that build up to a complete training loop, which is the thing PyTorch asks you to write yourself.

  1. 1

    Get comfortable with tensors

    Reshaping, views and adding dimensions — the source of most PyTorch error messages.

  2. 2

    Convert to and from NumPy

    Your data arrives as arrays and leaves as arrays, so this conversion is constant.

  3. 3

    Build a model from layers

    nn.Linear, activations, and stacking them into a network.

  4. 4

    Choose a loss and optimizer

    Cross-entropy for classification, MSE for regression, and Adam as a sensible default.

  5. 5

    Write the training loop

    DataLoader for batching, then forward, loss, backward, step — repeated.

Quick reference

The PyTorch calls you will use most

Twelve lines covering tensors, a model and the loop that trains it.

TaskCodeWorth knowing
Create a tensortorch.tensor([1.0, 2.0])Infers dtype. Pass dtype= to be sure.
From NumPytorch.from_numpy(arr)Shares memory with the array.
Back to NumPyt.numpy()Call .detach() first if it needs gradients.
Check the shapet.shapePrint it whenever an error mentions size.
Reshapet.view(-1, 784)reshape() is safer on non-contiguous tensors.
Add a batch dimensiont.unsqueeze(0)Models expect a batch, even of one.
Move to the GPUt.to(device)Model and data both, or you get a device error.
A simple modelnn.Sequential(nn.Linear(784, 10))Layers in order, like Keras.
Loss functionnn.CrossEntropyLoss()Takes raw logits — do not add softmax.
Optimizertorch.optim.Adam(model.parameters())Pass the parameters, not the model.
Clear gradientsopt.zero_grad()Skip this and gradients accumulate silently.
Backward and steploss.backward(); opt.step()Compute gradients, then update weights.

Every tutorial, by topic

Every PyTorch tutorial, grouped by what you are trying to do

Five topics instead of one flat list, following the order you meet them in when writing a training script.

Tensors

8

Reshaping with view and reshape, flattening, adding dimensions, stacking and concatenating, and converting to NumPy. Most PyTorch errors are shape errors, so this is the group to know well.

Layers and activations

11

The building blocks: linear and convolutional layers in one, two and three dimensions, batch normalisation, recurrent layers and activation functions such as leaky ReLU and tanh.

Building models

3

Assembling layers into a model, loading a saved one, printing a summary, and switching to evaluation mode with model.eval() before you predict.

Training

6

Loss functions for classification and regression, the Adam optimizer, DataLoader for batching, and stopping early when validation loss stops improving.

More PyTorch tutorials

4

MNIST end to end, resizing images, the JAX comparison and interview questions.

Keep going

What to learn next to it

Worth knowing what sits beside PyTorch before you commit to it.

Questions people ask

Frequently asked questions

Is PyTorch easier than TensorFlow?

Most people find it easier to debug, because it behaves like ordinary Python — you can print a tensor mid-model or step through the training loop. It asks you to write more code than Keras does, so it is more explicit rather than simpler.

What does optimizer.zero_grad() do and why is it needed?

PyTorch adds new gradients to whatever is already stored on each parameter rather than replacing them. That is deliberate and useful for techniques like gradient accumulation, but it means you must clear them before each batch or every step is computed from a mixture of batches.

What is the difference between view and reshape?

view returns a new view of the same memory and fails if the tensor is not contiguous. reshape does the same when it can and quietly copies when it cannot. Use reshape unless you specifically want the error.

Why do I get ‘Expected all tensors to be on the same device’?

Part of your computation is on the GPU and part on the CPU. Define device once, call model.to(device), and move every batch with xb.to(device) inside the loop. Tensors created fresh mid-loop are the usual culprits.

Do I add a softmax before CrossEntropyLoss?

No. nn.CrossEntropyLoss applies log-softmax internally, so it expects raw logits. Adding your own softmax first applies it twice and quietly weakens training. Use nn.NLLLoss if you really do want to pass log-probabilities.

Should I use PyTorch or Keras for a first deep learning project?

Keras if the goal is a working result quickly; PyTorch if you want to understand every step or plan to follow research papers. The concepts are the same either way, so switching later is a matter of days.

Write one training loop by hand

Forward, loss, zero_grad, backward, step. Once that loop makes sense, the rest of PyTorch is detail.