How to Find the First Number in a String in Python

To find the first number in a string in Python, use re.search(r"\d+", text). It returns the first run of digits, or None if the string has no digits, so always check the match before calling .group(). Below: a reusable function, versions without regex, how to handle negative numbers, decimals and thousands separators, and a few edge cases that trip up isdigit().

All examples were run with Python 3.12.5 and pandas 3.0.6; the output is copied from the terminal. For the full reference, see the re module documentation.

Quick answer

import re

text = "Order 4521 shipped in 3 boxes"

match = re.search(r"\d+", text)
if match:
    print(int(match.group()))

Output:

4521
Command Prompt screenshot: re.search finds 4521, the first number in a Python string
re.search() run in the Command Prompt.

A reusable first_number() function

Wrap the search in a function that returns None (or a default you choose) when there is no number, so your program never crashes with AttributeError: 'NoneType' object has no attribute 'group':

import re

def first_number(text, default=None):
    """Return the first whole number in text as an int, or default if there is none."""
    match = re.search(r"\d+", text)
    return int(match.group()) if match else default

print(first_number("Room 204, floor 2"))
print(first_number("no digits here"))
print(first_number("no digits here", default=0))

Output:

204
None
0

Find the first digit only

Use \d without the + to get one digit. Without regex, next() with a generator stops at the first digit it finds:

import re

text = "Invoice INV-7890"

# the first single digit
print(re.search(r"\d", text).group())

# the same without regex
print(next((ch for ch in text if ch.isdecimal()), None))

Output:

7
7

This finds the first digit in a string. For the first digit of an integer such as 7890, see get the first digit of a number in Python.

Without regex: a simple loop

Collect digits until the first non-digit after the number. This is easy to read and has no imports:

def first_number(text):
    """Find the first run of digits without the re module."""
    digits = ""
    for ch in text:
        if ch.isdecimal():
            digits += ch          # keep collecting digits
        elif digits:
            break                 # the first number has ended
    return int(digits) if digits else None

print(first_number("Flight BA2490 departs at 14:35"))
print(first_number("Gate B"))

Output:

2490
None
Command Prompt screenshot of a Python function that finds the first number in a string without regex
The loop version returns 2490, or None when there is no number.

Where is the first number? (index and span)

The match object also tells you where the number is, which is useful when you need the text before or after it:

import re

text = "Temperature today: 23 C, tomorrow: 19 C"

match = re.search(r"\d+", text)
print("number:", match.group())
print("starts at index:", match.start())
print("span:", match.span())
print("text before it:", repr(text[:match.start()]))

Output:

number: 23
starts at index: 19
span: (19, 21)
text before it: 'Temperature today: '

Negative numbers and decimals

\d+ only finds digits, so “-42.75” would give 42. Add an optional sign and decimal part to the pattern:

import re

PATTERN = r"[-+]?\d*\.?\d+"      # optional sign, optional decimals, allows ".5"

for text in ["Balance: -42.75 USD", "Growth +3.5%", "Rate .5 per cent", "Version 2"]:
    match = re.search(PATTERN, text)
    print(f"{text!r:24} -> {float(match.group())}")

Output:

'Balance: -42.75 USD'    -> -42.75
'Growth +3.5%'           -> 3.5
'Rate .5 per cent'       -> 0.5
'Version 2'              -> 2.0

Numbers with thousands separators

For amounts like 1,250,000.50, match the comma groups first, then remove the commas before converting:

import re

# 1,250,000.50 style numbers (commas as thousands separators) or plain numbers
PATTERN = r"-?\d{1,3}(?:,\d{3})+(?:\.\d+)?|-?\d+(?:\.\d+)?"

for text in ["Price: 1,250,000.50 GBP", "Only 42 left", "Loss of -3,400 this year"]:
    number = re.search(PATTERN, text).group()
    print(f"{number:>14} -> {float(number.replace(',', ''))}")

Output:

  1,250,000.50 -> 1250000.5
            42 -> 42.0
        -3,400 -> -3400.0
Command Prompt screenshot finding the first number with thousands separators and decimals in Python strings
Numbers like 1,250,000.50 and -3,400 are parsed correctly.
Python find first number in string: what re.search matches with the patterns d, d+, a signed decimal pattern and a thousands separator pattern on the same text
The same text, four patterns: the highlighted part is what re.search() returns.

Watch out: isdigit() accepts characters int() can’t convert

str.isdigit() is True for superscripts like ², but int() cannot convert them. isdecimal() only accepts real decimal digits, which is why the loop examples above use it:

for ch in "5²3":
    print(ch, "isdigit:", ch.isdigit(), " isdecimal:", ch.isdecimal())

try:
    print(int("²"))
except ValueError as err:
    print("ValueError:", err)

Output:

5 isdigit: True  isdecimal: True
² isdigit: True  isdecimal: False
3 isdigit: True  isdecimal: True
ValueError: invalid literal for int() with base 10: '²'

\d in re matches decimal digits from any script, and int() converts them. Add re.ASCII (or use [0-9]) if you only want 0 to 9:

import re

text = "Total: ٣٤ items"          # Arabic-Indic digits for 34

print(re.search(r"\d+", text).group())                 # \d matches any Unicode digit
print(int(re.search(r"\d+", text).group()))           # and int() understands them
print(re.search(r"\d+", text, re.ASCII))              # ASCII only: 0-9

Output:

٣٤
34
None

How the methods compare on tricky strings

Each cell is the real result of that method on the string (“loop” is the isdecimal() loop above):

Text\d+signed / decimalthousandsloop
'abc123def45''123''123''123''123'
'Price: $19.99''19''19.99''19.99''19'
'Temp -5 C''5''-5''-5''5'
'1,500 visitors''1''1''1,500''1'
'no digits'NoneNoneNoneNone
'x²'NoneNoneNoneNone
'Room Ù£''Ù£''Ù£''Ù£''Ù£'

Pick the pattern that matches your data: \d+ for IDs and counts, the signed/decimal pattern for prices and measurements, and the thousands pattern for formatted amounts.

First number from every string in a list

The walrus operator (:=, Python 3.8+) keeps the list comprehension to one line and still handles strings without a number:

import re

lines = ["ID 17 - Alice", "Bob (no id)", "ID 203 - Carol", "ID 9 - Dan"]

numbers = [int(m.group()) if (m := re.search(r"\d+", s)) else None for s in lines]
print(numbers)

Output:

[17, None, 203, 9]

First number in a pandas column

Series.str.extract() returns the first match in each row. The nullable Int64 type keeps rows without a number as <NA>:

import pandas as pd

df = pd.DataFrame({"product": ["Laptop 15 inch", "Phone X", "Monitor 27in 144Hz", "Cable 2m"]})

# extract() returns the first match of the group in each row (NaN if none)
df["first_number"] = df["product"].str.extract(r"(\d+)", expand=False).astype("Int64")
print(df)

Output:

              product  first_number
0      Laptop 15 inch            15
1             Phone X          <NA>
2  Monitor 27in 144Hz            27
3            Cable 2m             2

Is regex or the loop faster?

Both are fast enough for almost any program. On this machine the compiled regex was faster:

import re
import timeit

text = "Customer reference: ABCD-EFGH / order 99821 / 3 items"
pattern = re.compile(r"\d+")

def with_regex():
    m = pattern.search(text)
    return int(m.group()) if m else None

def with_loop():
    digits = ""
    for ch in text:
        if ch.isdecimal():
            digits += ch
        elif digits:
            break
    return int(digits) if digits else None

assert with_regex() == with_loop() == 99821
for func in (with_regex, with_loop):
    seconds = timeit.timeit(func, number=200_000)
    print(f"{func.__name__:11} {seconds * 1e6 / 200_000:.2f} microseconds per call")

Output:

with_regex  0.80 microseconds per call
with_loop   1.49 microseconds per call

Timings vary by computer and string length. Pick the version you find easier to read; switch only if profiling shows it matters.

Related tasks

More string tutorials you may find useful:

Frequently asked questions

How do I get the first number from a string in Python?

Use re.search(r"\d+", text) and convert the match: int(match.group()). Check that the match is not None first.

How do I find the first digit in a string?

Use re.search(r"\d", text), or without regex next((ch for ch in text if ch.isdecimal()), None).

How do I extract the first decimal or negative number?

Use a pattern with an optional sign and decimal part, such as r"[-+]?\d*\.?\d+", and convert with float().

What does re.search return if there is no number?

None. Calling .group() on it raises AttributeError, so test the result with if match: first.

Should I use isdigit() or isdecimal()?

isdecimal(). isdigit() also accepts characters such as superscript ² that int() cannot convert.