nn.Conv1d is the 1D convolution layer in PyTorch, for signals, time series and text. It expects a 3D tensor of (batch, channels, length):
conv = nn.Conv1d(in_channels=3, out_channels=16, kernel_size=5)
y = conv(torch.randn(8, 3, 100)) # (8, 16, 96)
Nearly every problem with it is a shape problem, and the two shapes to keep straight are that input and the weight, which is (out_channels, in_channels, kernel_size).
Runs below are on PyTorch 2.14.0, Python 3.12.5.
Input and output shapes of Conv1d
Channels sit in the middle, not last. That’s the opposite of how most datasets arrive, so you’ll be reshaping more often than not:
import torch
import torch.nn as nn
conv = nn.Conv1d(in_channels=3, out_channels=16, kernel_size=5)
x = torch.randn(8, 3, 100) # (batch, channels, length)
y = conv(x)
print("input :", tuple(x.shape))
print("output :", tuple(y.shape))
print("weight :", tuple(conv.weight.shape), "<- (out_channels, in_channels, kernel_size)")
print("bias :", tuple(conv.bias.shape))
print("params :", sum(p.numel() for p in conv.parameters()))
Output:
input : (8, 3, 100)
output : (8, 16, 96)
weight : (16, 3, 5) <- (out_channels, in_channels, kernel_size)
bias : (16,)
params : 256
The length shrank by 4 because a kernel of 5 loses kernel_size - 1 samples when there’s no padding. The channel count simply becomes out_channels.
The Conv1d weight shape
This is the thing most people come here to check. The weight is always three dimensions, in this order:
conv.weight.shape == (out_channels, in_channels, kernel_size)
import torch.nn as nn
for in_c, out_c, k in ((1, 8, 3), (3, 16, 5), (64, 32, 7)):
conv = nn.Conv1d(in_c, out_c, k)
weights = conv.weight.numel()
print(f"Conv1d({in_c}, {out_c}, {k}) -> weight {tuple(conv.weight.shape)}"
f" = {out_c}x{in_c}x{k} = {weights} + {out_c} bias")
Output:
Conv1d(1, 8, 3) -> weight (8, 1, 3) = 8x1x3 = 24 + 8 bias
Conv1d(3, 16, 5) -> weight (16, 3, 5) = 16x3x5 = 240 + 16 bias
Conv1d(64, 32, 7) -> weight (32, 64, 7) = 32x64x7 = 14336 + 32 bias
So the parameter count is out_channels × in_channels × kernel_size, plus one bias per output channel.
It reads in the same order as the 2D version, which has one extra kernel dimension on the end.
How do you calculate the Conv1d output length?
The formula is the same one Conv2d uses, applied to a single axis:
out = (length + 2*padding - dilation*(kernel-1) - 1) // stride + 1
import torch
import torch.nn as nn
def out_length(length, kernel, stride=1, padding=0, dilation=1):
return (length + 2 * padding - dilation * (kernel - 1) - 1) // stride + 1
x = torch.randn(1, 1, 100)
for kernel, stride, padding in ((3, 1, 0), (5, 1, 2), (3, 2, 1), (3, 1, 0)):
conv = nn.Conv1d(1, 1, kernel, stride=stride, padding=padding)
real = conv(x).shape[-1]
calc = out_length(100, kernel, stride, padding)
print(f"k={kernel} s={stride} p={padding} -> layer {real:>3}, formula {calc:>3}, agree {real == calc}")
same = nn.Conv1d(1, 1, 5, padding="same")
print("padding='same' keeps the length:", same(x).shape[-1])
Output:
k=3 s=1 p=0 -> layer 98, formula 98, agree True
k=5 s=1 p=2 -> layer 100, formula 100, agree True
k=3 s=2 p=1 -> layer 50, formula 50, agree True
k=3 s=1 p=0 -> layer 98, formula 98, agree True
padding='same' keeps the length: 100
padding="same" saves you the arithmetic when you want the length preserved, and it’s the usual choice in a text model where you don’t want to lose tokens at the edges.
Why Conv1d says you have the wrong number of channels
Feed Conv1d a batch of raw sequences shaped (batch, length) and you get an error about channels, which reads like nonsense until you know why:
import torch
import torch.nn as nn
conv = nn.Conv1d(1, 4, 3)
x = torch.randn(32, 100) # (batch, length) -- the channel dimension is missing
conv(x)
What PyTorch raises:
Traceback (most recent call last):
File "C:\pyguides\conv1d_shape_error.py", line 7, in <module>
conv(x)
File "C:\pyguides\venv\Lib\site-packages\torch\nn\modules\module.py", line 1783, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\pyguides\venv\Lib\site-packages\torch\nn\modules\module.py", line 1794, in _call_impl
return forward_call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\pyguides\venv\Lib\site-packages\torch\nn\modules\conv.py", line 385, in forward
return self._conv_forward(input, self.weight, self.bias)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\pyguides\venv\Lib\site-packages\torch\nn\modules\conv.py", line 380, in _conv_forward
return F.conv1d(
^^^^^^^^^
RuntimeError: Given groups=1, weight of size [4, 1, 3], expected input[1, 32, 100] to have 1 channels, but got 32 channels instead
Conv1d accepts unbatched input, so a 2D tensor is read as (channels, length). Your 32 sequences were taken as 32 channels of a single sequence, and the layer only wanted 1.
That is why the message talks about channels rather than dimensions. The fix is one call — unsqueeze(1) puts the channel axis where it belongs:
import torch
import torch.nn as nn
conv = nn.Conv1d(1, 4, 3)
x = torch.randn(32, 100) # (batch, length)
print("wrong shape:", tuple(x.shape))
fixed = x.unsqueeze(1) # add the channel dimension
print("fixed shape:", tuple(fixed.shape), "-> output", tuple(conv(fixed).shape))
# a single unbatched sequence is allowed too: (channels, length)
single = torch.randn(1, 100)
print("unbatched :", tuple(single.shape), "-> output", tuple(conv(single).shape))
Output:
wrong shape: (32, 100)
fixed shape: (32, 1, 100) -> output (32, 4, 98)
unbatched : (1, 100) -> output (4, 98)
The last line shows the flip side: a genuine unbatched (1, 100) works fine and returns (4, 98) with no batch axis. Useful, and exactly what makes the earlier error so confusing.
Conv1d on text: the transpose everyone forgets
nn.Embedding gives you (batch, length, channels). Conv1d wants (batch, channels, length). They are not the same and nothing warns you:
import torch
import torch.nn as nn
batch, seq_len, vocab, embed = 4, 50, 1000, 64
tokens = torch.randint(0, vocab, (batch, seq_len))
embedding = nn.Embedding(vocab, embed)
sequence = embedding(tokens)
print("after embedding:", tuple(sequence.shape), "<- (batch, length, channels)")
# Conv1d wants channels in the middle, so the axes have to be swapped
sequence = sequence.transpose(1, 2)
print("after transpose:", tuple(sequence.shape), "<- (batch, channels, length)")
conv = nn.Conv1d(embed, 128, kernel_size=3, padding=1)
print("after conv1d :", tuple(conv(sequence).shape))
pooled = conv(sequence).max(dim=2).values
print("after max pool :", tuple(pooled.shape), "<- one vector per sequence")
Output:
after embedding: (4, 50, 64) <- (batch, length, channels)
after transpose: (4, 64, 50) <- (batch, channels, length)
after conv1d : (4, 128, 50)
after max pool : (4, 128) <- one vector per sequence
Without the transpose the layer treats sequence positions as channels. It still runs and still trains, which is why this one survives so long before anybody notices.
Max pooling over the length axis at the end gives you one vector per sequence, ready for a classifier. That’s the whole shape of a text CNN.
Conv1d vs Conv2d
The difference is how many axes the kernel slides along, not how many dimensions the data has:
import torch
import torch.nn as nn
signal = torch.randn(1, 3, 64) # 3 channels, 64 samples
image = torch.randn(1, 3, 64, 64) # 3 channels, 64x64 pixels
c1 = nn.Conv1d(3, 8, 3)
c2 = nn.Conv2d(3, 8, 3)
print("Conv1d weight:", tuple(c1.weight.shape), "-> slides along ONE axis")
print("Conv2d weight:", tuple(c2.weight.shape), "-> slides along TWO axes")
print("Conv1d out :", tuple(c1(signal).shape))
print("Conv2d out :", tuple(c2(image).shape))
print("params 1d/2d :", c1.weight.numel(), "/", c2.weight.numel())
Output:
Conv1d weight: (8, 3, 3) -> slides along ONE axis
Conv2d weight: (8, 3, 3, 3) -> slides along TWO axes
Conv1d out : (1, 8, 62)
Conv2d out : (1, 8, 62, 62)
params 1d/2d : 72 / 216
nn.Conv1d | nn.Conv2d | |
|---|---|---|
| Input | (N, C, L) | (N, C, H, W) |
| Weight | (out, in, k) | (out, in, kH, kW) |
| Slides along | One axis | Two axes |
| Typical data | Audio, sensors, text | Images |
A 1D convolution over an image would be unusual but legal. What matters is which axis carries the ordering you want the kernel to respect, and the rest of the arguments behave exactly as in nn.Conv2d.
More PyTorch guides on this site:
- PyTorch nn.Conv2d
- PyTorch nn.Linear
- PyTorch MSELoss
- torch.cat for joining tensors
- torch.stack in PyTorch
Frequently asked questions
What input shape does nn.Conv1d expect?
(batch, channels, length), or (channels, length) without a batch. The signature is in the nn.Conv1d reference.
What is the weight shape of Conv1d?
(out_channels, in_channels, kernel_size), so the parameter count is the product of those three plus one bias per output channel.
How do I calculate the Conv1d output length?
(length + 2*padding - dilation*(kernel-1) - 1) // stride + 1. Use padding='same' to keep the length unchanged.
Why does Conv1d say my input has the wrong number of channels?
A 2D tensor is read as unbatched (channels, length), so (32, 100) becomes 32 channels. Add the channel axis with x.unsqueeze(1).
Why do I need to transpose after nn.Embedding?
Embedding returns (batch, length, channels) but Conv1d wants channels in the middle. Use .transpose(1, 2).
What is the difference between Conv1d and Conv2d?
Conv1d slides its kernel along one axis and Conv2d along two. The weight gains an extra kernel dimension in the 2D case.
Can I use Conv1d for time series?
Yes, that is its main use outside text. Put each measured variable in a channel and time along the length axis.
Bijay Kumar is a 13-time Microsoft MVP with more than 18 years in software development, and the founder of Python Guides and TSinfo Technologies. He started out building .NET and SharePoint solutions at HP, TCS and KPIT before moving into Python, machine learning and AI, and he also builds web apps with TypeScript and React. He writes the tutorials here himself, and every example is run before publishing so you see the real output. More about Bijay · Microsoft MVP profile · LinkedIn