Split a String Every N Characters in Python (Chunks)

Splitting a Python string every n characters is a one-line comprehension built on slicing:

text = "abcdefghij"
[text[i:i + 3] for i in range(0, len(text), 3)]
# ['abc', 'def', 'ghi', 'j']

The last piece comes out short rather than padded, which is usually what you want. There is no built-in string method for this, so slicing is the standard answer.

Python 3.12 added itertools.batched as an alternative, and two popular suggestions are worth a warning. Output comes from Python 3.12.5.

Diagram showing a Python string split into three-character chunks with a shorter final chunk
Ten characters split every three, leaving a short chunk at the end.

Split a string every n characters with slicing

The comprehension does two things at once, so it is worth reading slowly:

text = "abcdefghij"
n = 3

chunks = [text[i:i + n] for i in range(0, len(text), n)]

print("text   ->", text)
print("chunks ->", chunks)
print()
print("range(0, len(text), 3) gives the starting positions 0, 3, 6, 9")
print("and each slice takes the next three characters from there.")
print()
print("the last chunk is short rather than padded:")
print("  len of each ->", [len(c) for c in chunks])

Output:

text   -> abcdefghij
chunks -> ['abc', 'def', 'ghi', 'j']

range(0, len(text), 3) gives the starting positions 0, 3, 6, 9
and each slice takes the next three characters from there.

the last chunk is short rather than padded:
  len of each -> [3, 3, 3, 1]
Command Prompt showing a Python list comprehension splitting a string into chunks of three characters
range(0, len(text), n) gives the start of each chunk.

range(0, len(text), 3) produces 0, 3, 6, 9. Each slice then takes three characters from that position, and Python stops slicing at the end of the string rather than raising.

That last detail is what makes the short final chunk work without any special case. How range works covers the step argument.

Choosing the chunk size

The same line handles any size, including awkward ones:

text = "abcdefghij"

for n in (2, 4, 5, 20):
    chunks = [text[i:i + n] for i in range(0, len(text), n)]
    print(f"  every {n:2} characters -> {chunks}")

print()
print("a size larger than the string gives one chunk holding all of it.")
print()

print("an empty string gives no chunks at all:")
print("  ", [""[i:i + 3] for i in range(0, len(""), 3)])
print()

print("a size of zero is an error, because the loop would never advance:")
try:
    [text[i:i + 0] for i in range(0, len(text), 0)]
except ValueError as err:
    print("  ValueError:", err)

Output:

  every  2 characters -> ['ab', 'cd', 'ef', 'gh', 'ij']
  every  4 characters -> ['abcd', 'efgh', 'ij']
  every  5 characters -> ['abcde', 'fghij']
  every 20 characters -> ['abcdefghij']

a size larger than the string gives one chunk holding all of it.

an empty string gives no chunks at all:
   []

a size of zero is an error, because the loop would never advance:
  ValueError: range() arg 3 must not be zero
Command Prompt showing a Python string split every 2 4 5 and 20 characters
Any size works, and a size larger than the string gives one chunk.

A size of zero raises ValueError: range() arg 3 must not be zero, because the loop would never move forward. An empty string gives an empty list, with no error.

Chunks of n versus every nth character

These sound the same in English and do completely different things:

text = "abcdefghij"

print("two different questions that look alike:")
print()
print("chunks of 3 characters:")
print("  [text[i:i+3] for i in range(0, len(text), 3)]")
print("  ->", [text[i:i + 3] for i in range(0, len(text), 3)])
print()
print("every 3rd character:")
print("  text[::3]")
print("  ->", text[::3])
print()
print("the first keeps everything, grouped. The second throws away")
print("two characters out of every three.")

Output:

two different questions that look alike:

chunks of 3 characters:
  [text[i:i+3] for i in range(0, len(text), 3)]
  -> ['abc', 'def', 'ghi', 'j']

every 3rd character:
  text[::3]
  -> adgj

the first keeps everything, grouped. The second throws away
two characters out of every three.
Command Prompt comparing chunks of three characters with taking every third character in Python
Chunking keeps everything; a step slice discards what it skips.

text[::3] keeps the 1st, 4th and 7th characters and throws the rest away. The comprehension keeps every character and only groups them.

If you meant “take every third letter”, the step slice is right. If you meant “break this into threes”, it is not.

Using itertools.batched in Python 3.12

The standard library finally grew a function for this:

from itertools import batched

text = "abcdefghij"

print("batched() arrived in Python 3.12:")
print("  list(batched(text, 3)) ->", list(batched(text, 3)))
print()
print("it yields tuples, so join them back into strings:")
print("  ", ["".join(piece) for piece in batched(text, 3)])
print()
print("it works on any iterable, not just strings:")
print("  ", [list(b) for b in batched(range(10), 4)])
print()
print("the last batch is short, the same as the slicing version.")

Output:

batched() arrived in Python 3.12:
  list(batched(text, 3)) -> [('a', 'b', 'c'), ('d', 'e', 'f'), ('g', 'h', 'i'), ('j',)]

it yields tuples, so join them back into strings:
   ['abc', 'def', 'ghi', 'j']

it works on any iterable, not just strings:
   [[0, 1, 2, 3], [4, 5, 6, 7], [8, 9]]

the last batch is short, the same as the slicing version.

batched() yields tuples rather than strings, so joining them is usually the next step. It works on any iterable, which slicing does not, so it also handles generators and files.

Slicing remains the right choice on Python 3.11 and earlier, and it is no slower.

Why textwrap.wrap is not the answer

This one is suggested constantly, and it quietly does something else:

import textwrap

text = "abcdefghij"
print("on a string with no spaces, textwrap looks correct:")
print("  textwrap.wrap(text, 3) ->", textwrap.wrap(text, 3))
print()

sentence = "hello world foo"
print("but on real text it breaks at WORDS, not at a character count:")
print("  textwrap.wrap(sentence, 5) ->", textwrap.wrap(sentence, 5))
print("  expected 5-character chunks, got whole words")
print()
print("compare the slicing version on the same input:")
n = 5
print("  ", [sentence[i:i + n] for i in range(0, len(sentence), n)])
print()
print("textwrap is for laying out paragraphs. Use slicing for")
print("fixed-width chunks.")

Output:

on a string with no spaces, textwrap looks correct:
  textwrap.wrap(text, 3) -> ['abc', 'def', 'ghi', 'j']

but on real text it breaks at WORDS, not at a character count:
  textwrap.wrap(sentence, 5) -> ['hello', 'world', 'foo']
  expected 5-character chunks, got whole words

compare the slicing version on the same input:
   ['hello', ' worl', 'd foo']

textwrap is for laying out paragraphs. Use slicing for
fixed-width chunks.
Command Prompt showing that Python textwrap.wrap breaks at word boundaries instead of at a character count
textwrap.wrap() breaks at words, so chunks are not a fixed width.

On a string with no spaces it happens to look correct, which is how it ends up recommended. Give it real text and it breaks at word boundaries instead, producing pieces of whatever length the words happen to be.

It is a paragraph layout tool. For fixed-width chunks it is the wrong instrument, however convincing the first example looks.

Splitting with a regular expression

re.findall works, with two traps:

import re

text = "abcdefghij"

print("re.findall with a repetition count:")
print("  re.findall('.{1,3}', text) ->", re.findall(".{1,3}", text))
print()
print("without the 1, the remainder is silently DROPPED:")
print("  re.findall('.{3}', text)   ->", re.findall(".{3}", text))
print("  the trailing 'j' is gone")
print()

multiline = "ab\ncd"
print("and a dot does not match a newline by default:")
print("  re.findall('.{1,3}', 'ab\\ncd')        ->", re.findall(".{1,3}", multiline))
print("  re.findall('.{1,3}', 'ab\\ncd', re.S)  ->", re.findall(".{1,3}", multiline, re.S))
print("  the first version lost the newline entirely")
print()
print("slicing has neither problem, which is why it stays the default.")

Output:

re.findall with a repetition count:
  re.findall('.{1,3}', text) -> ['abc', 'def', 'ghi', 'j']

without the 1, the remainder is silently DROPPED:
  re.findall('.{3}', text)   -> ['abc', 'def', 'ghi']
  the trailing 'j' is gone

and a dot does not match a newline by default:
  re.findall('.{1,3}', 'ab\ncd')        -> ['ab', 'cd']
  re.findall('.{1,3}', 'ab\ncd', re.S)  -> ['ab\n', 'cd']
  the first version lost the newline entirely

slicing has neither problem, which is why it stays the default.
Command Prompt showing that a Python regex with an exact repetition count drops the remainder of the string
.{3} drops the remainder, and a dot skips newlines.

.{3} matches exactly three characters, so anything left over at the end is silently discarded. Writing .{1,3} keeps it.

A dot also refuses to match a newline unless you pass re.S, so multi-line text loses its line breaks. Slicing has neither problem. Splitting strings with regex covers the cases where a pattern genuinely helps.

A generator for large strings

When the input is big, produce chunks one at a time:

def chunks(text, n):
    """Yield successive n-character pieces of text."""
    for i in range(0, len(text), n):
        yield text[i:i + n]

sample = "abcdefghij"

print("as a list:")
print(" ", list(chunks(sample, 4)))
print()
print("one at a time, which never builds the whole list:")
for piece in chunks(sample, 4):
    print("  ", piece)
print()
print("on a large file this matters. The list version holds every")
print("chunk in memory at once; the generator holds one.")

Output:

as a list:
  ['abcd', 'efgh', 'ij']

one at a time, which never builds the whole list:
   abcd
   efgh
   ij

on a large file this matters. The list version holds every
chunk in memory at once; the generator holds one.

The comprehension builds every chunk before you use any of them. A generator holds one at a time, which matters on a file of any size.

Swapping between the two is a one-word change, so start with the comprehension and switch if memory becomes a problem.

Splitting a list into chunks of n

Exactly the same pattern, with a list in place of the string:

numbers = list(range(10))
n = 3

print("the same slice pattern works on a list:")
print(" ", [numbers[i:i + n] for i in range(0, len(numbers), n)])
print()

names = ["ann", "bob", "cat", "dan", "eve"]
print("and on a list of strings:")
print(" ", [names[i:i + 2] for i in range(0, len(names), 2)])
print()
print("only the thing being sliced changes; the loop is identical.")

Output:

the same slice pattern works on a list:
  [[0, 1, 2], [3, 4, 5], [6, 7, 8], [9]]

and on a list of strings:
  [['ann', 'bob'], ['cat', 'dan'], ['eve']]

only the thing being sliced changes; the loop is identical.

Slicing behaves identically on lists, tuples and strings, so nothing else changes. There is a dedicated guide at splitting a list into evenly sized chunks.

Padding or dropping the last chunk

The short final piece is easy to deal with either way:

text = "abcdefghij"
n = 4

chunks = [text[i:i + n] for i in range(0, len(text), n)]
print("raw chunks    ->", chunks)
print("lengths       ->", [len(c) for c in chunks])
print()

print("pad the short one to a fixed width:")
print("  ", [c.ljust(n, "_") for c in chunks])
print()

print("or drop it if a partial chunk is no use:")
print("  ", [c for c in chunks if len(c) == n])
print()

print("joining the unpadded chunks rebuilds the original exactly:")
print("  ", "".join(chunks) == text)

Output:

raw chunks    -> ['abcd', 'efgh', 'ij']
lengths       -> [4, 4, 2]

pad the short one to a fixed width:
   ['abcd', 'efgh', 'ij__']

or drop it if a partial chunk is no use:
   ['abcd', 'efgh']

joining the unpadded chunks rebuilds the original exactly:
   True

ljust(n, "_") pads it to a fixed width, which suits fixed-format output. Filtering on len(c) == n drops it when a partial chunk is meaningless.

Left alone, joining the chunks back together reproduces the original string exactly, which is a useful check.

Which method should you use?

SituationUse
Any version, any string[s[i:i+n] for i in range(0, len(s), n)]
Python 3.12 or later["".join(b) for b in batched(s, n)]
A very large string or a fileA generator with yield
A list instead of a stringThe same comprehension, unchanged
Every nth character, not chunkss[::n]
Laying out a paragraphtextwrap.wrap(), which breaks at words

Other string splitting guides worth a look:

Frequently asked questions

How do I split a string every n characters in Python?

Use a comprehension over a stepped range: [s[i:i+n] for i in range(0, len(s), n)]. The final chunk is shorter if the length does not divide evenly.

How do I split a string every 2 characters?

Set n to 2: [s[i:i+2] for i in range(0, len(s), 2)]. Nothing else about the pattern changes.

What is the difference between chunking and s[::n]?

Chunking groups every character into pieces of n. s[::n] keeps only every nth character and discards the rest.

Is there a built-in function for this?

itertools.batched(s, n) in Python 3.12 and later, which yields tuples you then join. Before that, slicing is the standard approach.

Why does textwrap.wrap give the wrong chunks?

Because it breaks at word boundaries, not at a character count. It looks correct on a string with no spaces, which is why it gets recommended.

How do I pad the last chunk to a full length?

Use ljust: [c.ljust(n, "_") for c in chunks]. Filter on len(c) == n instead if you would rather drop it.

Can I split a list the same way?

Yes, the comprehension is identical with a list in place of the string. Slicing behaviour is described in the Python sequence operations documentation.