Get the First N Characters of a String in Python

To get the first n characters of a string in Python, slice from the start:

text = "Hello, World"

text[:5]      # 'Hello'
text[:1]      # 'H'      the first character
text[:n]      # however many you need

The useful part is what it doesn’t do. Asking for more characters than the string has returns the whole string rather than raising an error.

All output below comes from real runs on Python 3.12.5.

Getting the first n characters of a string

The slice takes everything from the start up to, but not including, position n:

text = "Hello, World"

print(text[:5])          # first 5 characters
print(text[:1])          # first character
print(text[:10])         # first 10

n = 3
print(text[:n])          # however many you like

Output:

Hello
H
Hello, Wor
Hel
Command Prompt showing Python slice syntax returning the first five and first ten characters of a string
text[:n] with a literal number or a variable.

Leaving the first number out means “from the beginning”. text[:5] and text[0:5] are identical, and the shorter form is the usual one.

The count is a number of characters, so text[:10] gives ten characters, not position 10.

Why does a Python string slice never raise IndexError?

This is the behaviour worth knowing, because it removes a whole class of length checks:

text = "Hello, World"     # 12 characters

# asking for more than there is does NOT raise
print("text[:100] ->", repr(text[:100]))
print("text[:0]   ->", repr(text[:0]))

print()
# indexing does raise, slicing never does
empty = ""
try:
    print(empty[0])
except IndexError as err:
    print("empty[0]   -> IndexError:", err)

print("empty[:1]  ->", repr(empty[:1]), "  safe")

print()
print("that is why text[:1] is safer than text[0] for a first character")

Output:

text[:100] -> 'Hello, World'
text[:0]   -> ''

empty[0]   -> IndexError: string index out of range
empty[:1]  -> ''   safe

that is why text[:1] is safer than text[0] for a first character
Command Prompt showing that a Python slice beyond the string length returns the whole string while indexing raises IndexError
text[:100] is fine. empty[0] is not.
ExpressionOn a 12-character stringOn an empty string
text[:5]First 5 characters''
text[:100]The whole string''
text[0]First characterIndexError
text[:1]First character''

So text[:1] is the safe way to take a first character. text[0] crashes on an empty string, and empty strings turn up more often than you’d like.

If you do need to know the length first, that’s a separate check. See checking length in Python.

First character, last characters and everything between

The same syntax covers every variation people search for:

text = "Python programming is fun"

print("first 6        :", repr(text[:6]))
print("first character:", repr(text[:1]))
print("all but last 3 :", repr(text[:-3]))
print("last 3         :", repr(text[-3:]))
print("characters 7-18:", repr(text[7:18]))

print()
# every other character of the first ten
print("first 10, step 2:", repr(text[:10:2]))

Output:

first 6        : 'Python'
first character: 'P'
all but last 3 : 'Python programming is '
last 3         : 'fun'
characters 7-18: 'programming'

first 10, step 2: 'Pto r'
  • text[:1] — the first character, as a string.
  • text[-3:] — the last three characters.
  • text[:-3] — everything except the last three.
  • text[7:18] — a slice from the middle.
  • text[:10:2] — every other character of the first ten.

A negative number counts from the end. That’s how you remove the last character without knowing the length.

First n characters vs first n words

Slicing counts characters, so it will happily stop halfway through a word:

sentence = "Python string slicing is straightforward"

# first N characters is not the same as first N words
print("first 13 characters:", repr(sentence[:13]))
print("first 2 words      :", repr(" ".join(sentence.split()[:2])))

print()
# cutting mid-word, and cutting at a word boundary instead
cut = sentence[:20]
print("naive cut  :", repr(cut))
print("tidy cut   :", repr(cut.rsplit(" ", 1)[0]))

Output:

first 13 characters: 'Python string'
first 2 words      : 'Python string'

naive cut  : 'Python string slicin'
tidy cut   : 'Python string'

split() then a slice gives you words instead. Joining them back with a space restores the sentence.

rsplit(" ", 1)[0] is the quick fix for a character slice that landed mid-word: it drops the partial word at the end.

Truncating a Python string with textwrap.shorten

For anything a reader will see, textwrap.shorten beats a manual slice:

import textwrap

long_text = "Python string slicing is straightforward once you know it"

# a naive truncation can end mid-word and overshoot your width
print("naive  :", long_text[:27] + "...")

# textwrap.shorten never breaks a word and counts the placeholder
print("shorten:", textwrap.shorten(long_text, width=30, placeholder="..."))

print()
print("shorten result length:", len(textwrap.shorten(long_text, width=30, placeholder="...")))

Output:

naive  : Python string slicing is st...
shorten: Python string slicing is...

shorten result length: 27

It never breaks a word, and it counts the placeholder against the width, so the result really is within the limit you asked for.

A naive text[:27] + "..." produces 30 characters when you asked for 27, and often ends mid-word.

Slicing strings with accents and emoji

Python slices by code point, and a code point is not always a whole character on screen:

import unicodedata

# a family emoji is several code points joined together
family = "\U0001F469‍\U0001F467"
print("family emoji : len =", len(family), "code points")
print("  family[:1] ->", ascii(family[:1]), " half the emoji")

# an accent can be a separate character from the letter
cafe = "café"                 # 'e' followed by a combining acute
print("\ncafe + accent: len =", len(cafe))
print("  cafe[:4]   ->", ascii(cafe[:4]), " the accent was left behind")

# normalising combines them into one character first
fixed = unicodedata.normalize("NFC", cafe)
print("\nafter NFC    : len =", len(fixed))
print("  fixed[:4]  ->", ascii(fixed[:4]), " correct")

Output:

family emoji : len = 3 code points
  family[:1] -> '\U0001f469'  half the emoji

cafe + accent: len = 5
  cafe[:4]   -> 'cafe'  the accent was left behind

after NFC    : len = 4
  fixed[:4]  -> 'caf\xe9'  correct
Command Prompt showing that slicing a Python string with a combining accent or an emoji splits the character apart
len("cafe" + accent) is 5, and slicing to 4 leaves the accent behind.

An emoji like a family is several code points joined by invisible characters, so [:1] returns a fragment rather than the picture.

unicodedata.normalize("NFC", text) combines letter-plus-accent pairs into single characters first, which fixes the accent case.

Emoji need more than normalising. If you’re truncating user-generated text, a library such as regex with grapheme-cluster matching is the correct tool.

Getting the first n characters in pandas

A whole column at once uses the .str accessor, which slices every value:

import pandas as pd

df = pd.DataFrame({"code": ["AB-1234", "CD-5678", "EF-9012"]})

# .str slices every value in the column
df["prefix"] = df["code"].str[:2]
df["number"] = df["code"].str[3:]

print(df)

print()
print("dtype of the slice:", df["prefix"].dtype)

Output:

      code prefix number
0  AB-1234     AB   1234
1  CD-5678     CD   5678
2  EF-9012     EF   9012

dtype of the slice: str
Command Prompt showing a pandas DataFrame column sliced with str to take the first two characters
df["code"].str[:2] slices every row in one go.

The syntax after .str is the same slicing you’d use on a single string, so the same rules about over-long slices apply.

Looping over rows to do this is much slower and entirely unnecessary.

Common first-n-characters mistakes

SymptomCauseFix
IndexErrorUsed text[0] on an empty stringUse text[:1]
One character shortThought n was a positiontext[:n] gives n characters
Word cut in halfSliced by charactersplit(), or textwrap.shorten
Accent in the wrong placeCombining character split offNormalise with NFC first
Result longer than the limitAdded ... after slicingtextwrap.shorten counts it

Other Python string slicing guides:

Frequently asked questions

How do I get the first n characters of a string in Python?

Slice it: text[:n]. Slicing is described in the Python sequence operations reference.

How do I get the first character of a string?

text[:1] is safest, because it returns an empty string rather than raising when the string is empty. text[0] also works when you know there is at least one character.

What happens if n is longer than the string?

You get the whole string. Slicing clamps to the available length and never raises an IndexError.

How do I get