PyTorch nn.Conv1d: Shapes, Weights and Examples

nn.Conv1d is the 1D convolution layer in PyTorch, for signals, time series and text. It expects a 3D tensor of (batch, channels, length):

conv = nn.Conv1d(in_channels=3, out_channels=16, kernel_size=5)
y = conv(torch.randn(8, 3, 100))     # (8, 16, 96)

Nearly every problem with it is a shape problem, and the two shapes to keep straight are that input and the weight, which is (out_channels, in_channels, kernel_size).

Runs below are on PyTorch 2.14.0, Python 3.12.5.

Input and output shapes of Conv1d

Channels sit in the middle, not last. That’s the opposite of how most datasets arrive, so you’ll be reshaping more often than not:

import torch
import torch.nn as nn

conv = nn.Conv1d(in_channels=3, out_channels=16, kernel_size=5)

x = torch.randn(8, 3, 100)          # (batch, channels, length)
y = conv(x)

print("input  :", tuple(x.shape))
print("output :", tuple(y.shape))
print("weight :", tuple(conv.weight.shape), "<- (out_channels, in_channels, kernel_size)")
print("bias   :", tuple(conv.bias.shape))
print("params :", sum(p.numel() for p in conv.parameters()))

Output:

input  : (8, 3, 100)
output : (8, 16, 96)
weight : (16, 3, 5) <- (out_channels, in_channels, kernel_size)
bias   : (16,)
params : 256
Command Prompt showing a PyTorch Conv1d layer turning an 8 by 3 by 100 tensor into 8 by 16 by 96, with its weight shape and parameter count
3 channels in, 16 out, and the length drops from 100 to 96.

The length shrank by 4 because a kernel of 5 loses kernel_size - 1 samples when there’s no padding. The channel count simply becomes out_channels.

The Conv1d weight shape

This is the thing most people come here to check. The weight is always three dimensions, in this order:

conv.weight.shape == (out_channels, in_channels, kernel_size)
import torch.nn as nn

for in_c, out_c, k in ((1, 8, 3), (3, 16, 5), (64, 32, 7)):
    conv = nn.Conv1d(in_c, out_c, k)
    weights = conv.weight.numel()
    print(f"Conv1d({in_c}, {out_c}, {k}) -> weight {tuple(conv.weight.shape)}"
          f"  = {out_c}x{in_c}x{k} = {weights} + {out_c} bias")

Output:

Conv1d(1, 8, 3) -> weight (8, 1, 3)  = 8x1x3 = 24 + 8 bias
Conv1d(3, 16, 5) -> weight (16, 3, 5)  = 16x3x5 = 240 + 16 bias
Conv1d(64, 32, 7) -> weight (32, 64, 7)  = 32x64x7 = 14336 + 32 bias
Command Prompt showing Conv1d weight shapes for three different channel and kernel combinations with the parameter arithmetic
Out channels first, then in channels, then the kernel length.

So the parameter count is out_channels × in_channels × kernel_size, plus one bias per output channel.

It reads in the same order as the 2D version, which has one extra kernel dimension on the end.

How do you calculate the Conv1d output length?

The formula is the same one Conv2d uses, applied to a single axis:

out = (length + 2*padding - dilation*(kernel-1) - 1) // stride + 1
import torch
import torch.nn as nn

def out_length(length, kernel, stride=1, padding=0, dilation=1):
    return (length + 2 * padding - dilation * (kernel - 1) - 1) // stride + 1

x = torch.randn(1, 1, 100)

for kernel, stride, padding in ((3, 1, 0), (5, 1, 2), (3, 2, 1), (3, 1, 0)):
    conv = nn.Conv1d(1, 1, kernel, stride=stride, padding=padding)
    real = conv(x).shape[-1]
    calc = out_length(100, kernel, stride, padding)
    print(f"k={kernel} s={stride} p={padding} -> layer {real:>3}, formula {calc:>3}, agree {real == calc}")

same = nn.Conv1d(1, 1, 5, padding="same")
print("padding='same' keeps the length:", same(x).shape[-1])

Output:

k=3 s=1 p=0 -> layer  98, formula  98, agree True
k=5 s=1 p=2 -> layer 100, formula 100, agree True
k=3 s=2 p=1 -> layer  50, formula  50, agree True
k=3 s=1 p=0 -> layer  98, formula  98, agree True
padding='same' keeps the length: 100

padding="same" saves you the arithmetic when you want the length preserved, and it’s the usual choice in a text model where you don’t want to lose tokens at the edges.

Why Conv1d says you have the wrong number of channels

Feed Conv1d a batch of raw sequences shaped (batch, length) and you get an error about channels, which reads like nonsense until you know why:

import torch
import torch.nn as nn

conv = nn.Conv1d(1, 4, 3)

x = torch.randn(32, 100)        # (batch, length) -- the channel dimension is missing
conv(x)

What PyTorch raises:

Traceback (most recent call last):
  File "C:\pyguides\conv1d_shape_error.py", line 7, in <module>
    conv(x)
  File "C:\pyguides\venv\Lib\site-packages\torch\nn\modules\module.py", line 1783, in _wrapped_call_impl
    return self._call_impl(*args, **kwargs)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "C:\pyguides\venv\Lib\site-packages\torch\nn\modules\module.py", line 1794, in _call_impl
    return forward_call(*args, **kwargs)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "C:\pyguides\venv\Lib\site-packages\torch\nn\modules\conv.py", line 385, in forward
    return self._conv_forward(input, self.weight, self.bias)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "C:\pyguides\venv\Lib\site-packages\torch\nn\modules\conv.py", line 380, in _conv_forward
    return F.conv1d(
           ^^^^^^^^^
RuntimeError: Given groups=1, weight of size [4, 1, 3], expected input[1, 32, 100] to have 1 channels, but got 32 channels instead
Command Prompt showing a PyTorch Conv1d RuntimeError saying the input was expected to have 1 channel but got 32
32 sequences of 100 became 32 channels of one sequence.

Conv1d accepts unbatched input, so a 2D tensor is read as (channels, length). Your 32 sequences were taken as 32 channels of a single sequence, and the layer only wanted 1.

That is why the message talks about channels rather than dimensions. The fix is one call — unsqueeze(1) puts the channel axis where it belongs:

import torch
import torch.nn as nn

conv = nn.Conv1d(1, 4, 3)

x = torch.randn(32, 100)                 # (batch, length)
print("wrong shape:", tuple(x.shape))

fixed = x.unsqueeze(1)                   # add the channel dimension
print("fixed shape:", tuple(fixed.shape), "-> output", tuple(conv(fixed).shape))

# a single unbatched sequence is allowed too: (channels, length)
single = torch.randn(1, 100)
print("unbatched  :", tuple(single.shape), "-> output", tuple(conv(single).shape))

Output:

wrong shape: (32, 100)
fixed shape: (32, 1, 100) -> output (32, 4, 98)
unbatched  : (1, 100) -> output (4, 98)

The last line shows the flip side: a genuine unbatched (1, 100) works fine and returns (4, 98) with no batch axis. Useful, and exactly what makes the earlier error so confusing.

Conv1d on text: the transpose everyone forgets

nn.Embedding gives you (batch, length, channels). Conv1d wants (batch, channels, length). They are not the same and nothing warns you:

import torch
import torch.nn as nn

batch, seq_len, vocab, embed = 4, 50, 1000, 64

tokens = torch.randint(0, vocab, (batch, seq_len))
embedding = nn.Embedding(vocab, embed)

sequence = embedding(tokens)
print("after embedding:", tuple(sequence.shape), "<- (batch, length, channels)")

# Conv1d wants channels in the middle, so the axes have to be swapped
sequence = sequence.transpose(1, 2)
print("after transpose:", tuple(sequence.shape), "<- (batch, channels, length)")

conv = nn.Conv1d(embed, 128, kernel_size=3, padding=1)
print("after conv1d   :", tuple(conv(sequence).shape))

pooled = conv(sequence).max(dim=2).values
print("after max pool :", tuple(pooled.shape), "<- one vector per sequence")

Output:

after embedding: (4, 50, 64) <- (batch, length, channels)
after transpose: (4, 64, 50) <- (batch, channels, length)
after conv1d   : (4, 128, 50)
after max pool : (4, 128) <- one vector per sequence
Command Prompt showing an embedding output transposed before being passed to a PyTorch Conv1d layer and then max pooled
Embed, transpose, convolve, pool: the four shapes in order.

Without the transpose the layer treats sequence positions as channels. It still runs and still trains, which is why this one survives so long before anybody notices.

Max pooling over the length axis at the end gives you one vector per sequence, ready for a classifier. That’s the whole shape of a text CNN.

Conv1d vs Conv2d

The difference is how many axes the kernel slides along, not how many dimensions the data has:

import torch
import torch.nn as nn

signal = torch.randn(1, 3, 64)          # 3 channels, 64 samples
image = torch.randn(1, 3, 64, 64)       # 3 channels, 64x64 pixels

c1 = nn.Conv1d(3, 8, 3)
c2 = nn.Conv2d(3, 8, 3)

print("Conv1d weight:", tuple(c1.weight.shape), "-> slides along ONE axis")
print("Conv2d weight:", tuple(c2.weight.shape), "-> slides along TWO axes")
print("Conv1d out   :", tuple(c1(signal).shape))
print("Conv2d out   :", tuple(c2(image).shape))
print("params 1d/2d :", c1.weight.numel(), "/", c2.weight.numel())

Output:

Conv1d weight: (8, 3, 3) -> slides along ONE axis
Conv2d weight: (8, 3, 3, 3) -> slides along TWO axes
Conv1d out   : (1, 8, 62)
Conv2d out   : (1, 8, 62, 62)
params 1d/2d : 72 / 216
nn.Conv1dnn.Conv2d
Input(N, C, L)(N, C, H, W)
Weight(out, in, k)(out, in, kH, kW)
Slides alongOne axisTwo axes
Typical dataAudio, sensors, textImages

A 1D convolution over an image would be unusual but legal. What matters is which axis carries the ordering you want the kernel to respect, and the rest of the arguments behave exactly as in nn.Conv2d.

More PyTorch guides on this site:

Frequently asked questions

What input shape does nn.Conv1d expect?

(batch, channels, length), or (channels, length) without a batch. The signature is in the nn.Conv1d reference.

What is the weight shape of Conv1d?

(out_channels, in_channels, kernel_size), so the parameter count is the product of those three plus one bias per output channel.

How do I calculate the Conv1d output length?

(length + 2*padding - dilation*(kernel-1) - 1) // stride + 1. Use padding='same' to keep the length unchanged.

Why does Conv1d say my input has the wrong number of channels?

A 2D tensor is read as unbatched (channels, length), so (32, 100) becomes 32 channels. Add the channel axis with x.unsqueeze(1).

Why do I need to transpose after nn.Embedding?

Embedding returns (batch, length, channels) but Conv1d wants channels in the middle. Use .transpose(1, 2).

What is the difference between Conv1d and Conv2d?

Conv1d slides its kernel along one axis and Conv2d along two. The weight gains an extra kernel dimension in the 2D case.

Can I use Conv1d for time series?

Yes, that is its main use outside text. Put each measured variable in a channel and time along the length axis.