To find the first number in a string in Python, use re.search(r"\d+", text). It returns the first run of digits, or None if the string has no digits, so always check the match before calling .group(). Below: a reusable function, versions without regex, how to handle negative numbers, decimals and thousands separators, and a few edge cases that trip up isdigit().
All examples were run with Python 3.12.5 and pandas 3.0.6; the output is copied from the terminal. For the full reference, see the re module documentation.
Quick answer
import re
text = "Order 4521 shipped in 3 boxes"
match = re.search(r"\d+", text)
if match:
print(int(match.group()))
Output:
4521
re.search() run in the Command Prompt.A reusable first_number() function
Wrap the search in a function that returns None (or a default you choose) when there is no number, so your program never crashes with AttributeError: 'NoneType' object has no attribute 'group':
import re
def first_number(text, default=None):
"""Return the first whole number in text as an int, or default if there is none."""
match = re.search(r"\d+", text)
return int(match.group()) if match else default
print(first_number("Room 204, floor 2"))
print(first_number("no digits here"))
print(first_number("no digits here", default=0))
Output:
204
None
0
Find the first digit only
Use \d without the + to get one digit. Without regex, next() with a generator stops at the first digit it finds:
import re
text = "Invoice INV-7890"
# the first single digit
print(re.search(r"\d", text).group())
# the same without regex
print(next((ch for ch in text if ch.isdecimal()), None))
Output:
7
7
This finds the first digit in a string. For the first digit of an integer such as 7890, see get the first digit of a number in Python.
Without regex: a simple loop
Collect digits until the first non-digit after the number. This is easy to read and has no imports:
def first_number(text):
"""Find the first run of digits without the re module."""
digits = ""
for ch in text:
if ch.isdecimal():
digits += ch # keep collecting digits
elif digits:
break # the first number has ended
return int(digits) if digits else None
print(first_number("Flight BA2490 departs at 14:35"))
print(first_number("Gate B"))
Output:
2490
None
Where is the first number? (index and span)
The match object also tells you where the number is, which is useful when you need the text before or after it:
import re
text = "Temperature today: 23 C, tomorrow: 19 C"
match = re.search(r"\d+", text)
print("number:", match.group())
print("starts at index:", match.start())
print("span:", match.span())
print("text before it:", repr(text[:match.start()]))
Output:
number: 23
starts at index: 19
span: (19, 21)
text before it: 'Temperature today: '
Negative numbers and decimals
\d+ only finds digits, so “-42.75” would give 42. Add an optional sign and decimal part to the pattern:
import re
PATTERN = r"[-+]?\d*\.?\d+" # optional sign, optional decimals, allows ".5"
for text in ["Balance: -42.75 USD", "Growth +3.5%", "Rate .5 per cent", "Version 2"]:
match = re.search(PATTERN, text)
print(f"{text!r:24} -> {float(match.group())}")
Output:
'Balance: -42.75 USD' -> -42.75
'Growth +3.5%' -> 3.5
'Rate .5 per cent' -> 0.5
'Version 2' -> 2.0
Numbers with thousands separators
For amounts like 1,250,000.50, match the comma groups first, then remove the commas before converting:
import re
# 1,250,000.50 style numbers (commas as thousands separators) or plain numbers
PATTERN = r"-?\d{1,3}(?:,\d{3})+(?:\.\d+)?|-?\d+(?:\.\d+)?"
for text in ["Price: 1,250,000.50 GBP", "Only 42 left", "Loss of -3,400 this year"]:
number = re.search(PATTERN, text).group()
print(f"{number:>14} -> {float(number.replace(',', ''))}")
Output:
1,250,000.50 -> 1250000.5
42 -> 42.0
-3,400 -> -3400.0
Watch out: isdigit() accepts characters int() can’t convert
str.isdigit() is True for superscripts like ², but int() cannot convert them. isdecimal() only accepts real decimal digits, which is why the loop examples above use it:
for ch in "5²3":
print(ch, "isdigit:", ch.isdigit(), " isdecimal:", ch.isdecimal())
try:
print(int("²"))
except ValueError as err:
print("ValueError:", err)
Output:
5 isdigit: True isdecimal: True
² isdigit: True isdecimal: False
3 isdigit: True isdecimal: True
ValueError: invalid literal for int() with base 10: '²'
\d in re matches decimal digits from any script, and int() converts them. Add re.ASCII (or use [0-9]) if you only want 0 to 9:
import re
text = "Total: ٣٤ items" # Arabic-Indic digits for 34
print(re.search(r"\d+", text).group()) # \d matches any Unicode digit
print(int(re.search(r"\d+", text).group())) # and int() understands them
print(re.search(r"\d+", text, re.ASCII)) # ASCII only: 0-9
Output:
٣٤
34
None
How the methods compare on tricky strings
Each cell is the real result of that method on the string (“loop” is the isdecimal() loop above):
| Text | \d+ | signed / decimal | thousands | loop |
|---|---|---|---|---|
'abc123def45' | '123' | '123' | '123' | '123' |
'Price: $19.99' | '19' | '19.99' | '19.99' | '19' |
'Temp -5 C' | '5' | '-5' | '-5' | '5' |
'1,500 visitors' | '1' | '1' | '1,500' | '1' |
'no digits' | None | None | None | None |
'x²' | None | None | None | None |
'Room Ù£' | 'Ù£' | 'Ù£' | 'Ù£' | 'Ù£' |
Pick the pattern that matches your data: \d+ for IDs and counts, the signed/decimal pattern for prices and measurements, and the thousands pattern for formatted amounts.
First number from every string in a list
The walrus operator (:=, Python 3.8+) keeps the list comprehension to one line and still handles strings without a number:
import re
lines = ["ID 17 - Alice", "Bob (no id)", "ID 203 - Carol", "ID 9 - Dan"]
numbers = [int(m.group()) if (m := re.search(r"\d+", s)) else None for s in lines]
print(numbers)
Output:
[17, None, 203, 9]
First number in a pandas column
Series.str.extract() returns the first match in each row. The nullable Int64 type keeps rows without a number as <NA>:
import pandas as pd
df = pd.DataFrame({"product": ["Laptop 15 inch", "Phone X", "Monitor 27in 144Hz", "Cable 2m"]})
# extract() returns the first match of the group in each row (NaN if none)
df["first_number"] = df["product"].str.extract(r"(\d+)", expand=False).astype("Int64")
print(df)
Output:
product first_number
0 Laptop 15 inch 15
1 Phone X <NA>
2 Monitor 27in 144Hz 27
3 Cable 2m 2
Is regex or the loop faster?
Both are fast enough for almost any program. On this machine the compiled regex was faster:
import re
import timeit
text = "Customer reference: ABCD-EFGH / order 99821 / 3 items"
pattern = re.compile(r"\d+")
def with_regex():
m = pattern.search(text)
return int(m.group()) if m else None
def with_loop():
digits = ""
for ch in text:
if ch.isdecimal():
digits += ch
elif digits:
break
return int(digits) if digits else None
assert with_regex() == with_loop() == 99821
for func in (with_regex, with_loop):
seconds = timeit.timeit(func, number=200_000)
print(f"{func.__name__:11} {seconds * 1e6 / 200_000:.2f} microseconds per call")
Output:
with_regex 0.80 microseconds per call
with_loop 1.49 microseconds per call
Timings vary by computer and string length. Pick the version you find easier to read; switch only if profiling shows it matters.
Related tasks
- Extract all numbers from a string (use
re.findall()) - Find the last number in a string
- Check if a string begins with a number
- Count the numbers in a string
More string tutorials you may find useful:
- Compare strings in Python
- Split a string at every Nth character
- Remove special characters from a string
- Split a string into an array
Frequently asked questions
How do I get the first number from a string in Python?
Use re.search(r"\d+", text) and convert the match: int(match.group()). Check that the match is not None first.
How do I find the first digit in a string?
Use re.search(r"\d", text), or without regex next((ch for ch in text if ch.isdecimal()), None).
How do I extract the first decimal or negative number?
Use a pattern with an optional sign and decimal part, such as r"[-+]?\d*\.?\d+", and convert with float().
What does re.search return if there is no number?
None. Calling .group() on it raises AttributeError, so test the result with if match: first.
Should I use isdigit() or isdecimal()?
isdecimal(). isdigit() also accepts characters such as superscript ² that int() cannot convert.
Bijay Kumar is a 13-time Microsoft MVP with more than 18 years in software development, and the founder of Python Guides and TSinfo Technologies. He started out building .NET and SharePoint solutions at HP, TCS and KPIT before moving into Python, machine learning and AI, and he also builds web apps with TypeScript and React. He writes the tutorials here himself, and every example is run before publishing so you see the real output. More about Bijay · Microsoft MVP profile · LinkedIn