To change the datatype (data type) of a column in pandas, use astype(): df["price"] = df["price"].astype(float). Pass a dictionary to change several column types at once, use pd.to_numeric() or pd.to_datetime() when the data is messy, and astype("category") to save memory. This guide shows each way to change a pandas column type, the errors you will meet (and how to fix them), and how missing values affect integer columns.
All examples were run with Python 3.12.5 and pandas 3.0.6 in the Windows Command Prompt. Reference: DataFrame.astype and pandas.to_numeric in the pandas documentation.
Change a column type with astype()
Check the current types with df.dtypes, then assign the converted column back:
import pandas as pd
df = pd.DataFrame({
"order_id": ["1001", "1002", "1003"],
"price": ["19.99", "5.50", "12.00"],
"qty": [2.0, 1.0, 3.0],
"order_date": ["2025-03-14", "2025-03-15", "2025-03-16"],
})
print(df.dtypes, "\n")
df["order_id"] = df["order_id"].astype(int) # text -> integer
df["price"] = df["price"].astype(float) # text -> float
df["qty"] = df["qty"].astype("int64") # float -> integer
print(df.dtypes)
print(round(df["price"].sum(), 2)) # numbers now add up
Output:
order_id str
price str
qty float64
order_date str
dtype: object
order_id int64
price float64
qty int64
order_date str
dtype: object
37.49
int64 and float64.In pandas 3, text columns show the dtype str (earlier versions showed object). astype() returns a new Series; nothing changes until you assign it.
Change multiple column types at once
import pandas as pd
df = pd.DataFrame({
"order_id": ["1001", "1002", "1003"],
"price": ["19.99", "5.50", "12.00"],
"qty": [2.0, 1.0, 3.0],
"order_date": ["2025-03-14", "2025-03-15", "2025-03-16"],
})
df = df.astype({"order_id": "int64", "price": "float64", "qty": "int32"}) # several columns at once
print(df.dtypes, "\n")
df2 = df.astype(str) # every column
print(df2.dtypes)
Output:
order_id int64
price float64
qty int32
order_date str
dtype: object
order_id str
price str
qty str
order_date str
dtype: object
Convert messy text to numbers with pd.to_numeric()
astype(float) fails on the first value it cannot parse. pd.to_numeric(..., errors="coerce") turns bad values into NaN instead, and a little cleaning first rescues values like "1,200":
import pandas as pd
df = pd.DataFrame({"amount": ["120", "85.5", "N/A", "1,200", " 42 "]})
try:
df["amount"].astype(float)
except ValueError as e:
print("astype(float) -> ValueError:", e)
df["coerced"] = pd.to_numeric(df["amount"], errors="coerce") # bad values -> NaN
df["cleaned"] = pd.to_numeric(df["amount"].str.replace(",", "").str.strip(), errors="coerce")
print(df)
print(df.dtypes)
Output:
astype(float) -> ValueError: could not convert string to float: 'N/A'
amount coerced cleaned
0 120 120.0 120.0
1 85.5 85.5 85.5
2 N/A NaN NaN
3 1,200 NaN 1200.0
4 42 42.0 42.0
amount str
coerced float64
cleaned float64
dtype: object
errors="coerce" keeps going; cleaning first saves more values.Integers with missing values: use Int64
A normal integer column cannot hold NaN. Use pandas’ nullable "Int64" type (capital I), or fill the gaps first:
import pandas as pd
s = pd.Series([3.0, None, 7.0])
try:
s.astype(int)
except Exception as e:
print(type(e).__name__ + ":", e)
print(s.astype("Int64")) # nullable integer keeps the missing value as <NA>
print(s.fillna(0).astype(int).tolist()) # or fill first
Output:
IntCastingNaNError: Cannot convert non-finite values (NA or inf) to integer.Replace or remove non-finite values or cast to an integer typethat supports these values (e.g. 'Int64')
0 3
1 <NA>
2 7
dtype: Int64
[3, 0, 7]
astype(int) fails on NaN; astype("Int64") keeps it as <NA>.Convert a column to datetime
Use pd.to_datetime() for dates stored as text or numbers. Give the format when you know it, and errors="coerce" for rows that do not match:
import pandas as pd
df = pd.DataFrame({"shipped": ["2025-03-14", "14/03/2025", "not shipped"],
"epoch_s": [1741910400, 1741996800, 1742083200]})
df["shipped_dt"] = pd.to_datetime(df["shipped"], format="%Y-%m-%d", errors="coerce")
df["from_epoch"] = pd.to_datetime(df["epoch_s"], unit="s")
print(df)
print(df.dtypes)
print(df["from_epoch"].dt.day_name().tolist())
Output:
shipped epoch_s shipped_dt from_epoch
0 2025-03-14 1741910400 2025-03-14 2025-03-14
1 14/03/2025 1741996800 NaT 2025-03-15
2 not shipped 1742083200 NaT 2025-03-16
shipped str
epoch_s int64
shipped_dt datetime64[us]
from_epoch datetime64[s]
dtype: object
['Friday', 'Saturday', 'Sunday']
For formatting dates back to text, see convert a date and time to a string.
Save memory with the category type
A column with a few repeated values (states, departments, status codes) takes far less memory as category:
import pandas as pd
import numpy as np
rng = np.random.default_rng(0)
df = pd.DataFrame({"state": rng.choice(["TX", "CA", "NY", "FL"], size=100_000)})
before = df["state"].memory_usage(deep=True)
df["state"] = df["state"].astype("category")
after = df["state"].memory_usage(deep=True)
print(df["state"].dtype, list(df["state"].cat.categories))
print(f"memory: {before:,} -> {after:,} bytes ({before / after:.0f}x smaller)")
Output:
category ['CA', 'FL', 'NY', 'TX']
memory: 5,100,132 -> 100,336 bytes (51x smaller)
Convert a column to boolean safely
astype(bool) treats every non-empty string as True, including "False". Map the values explicitly:
import pandas as pd
s = pd.Series(["True", "False", "false", "yes", ""])
print(s.astype(bool).tolist()) # any non-empty string is True!
mapping = {"true": True, "false": False, "yes": True, "no": False}
print(s.str.lower().map(mapping).tolist()) # explicit mapping
Output:
[True, True, True, True, False]
[True, False, False, True, nan]
Let pandas choose: convert_dtypes()
convert_dtypes() picks the best nullable type for every column, which is handy right after loading a file:
import pandas as pd
df = pd.DataFrame({"id": [1, 2, None], "name": ["Ana", "Ben", None], "score": [9.5, None, 7.0], "active": [True, None, False]})
print(df.dtypes, "\n")
print(df.convert_dtypes().dtypes) # best nullable dtype for every column
Output:
id float64
name str
score float64
active object
dtype: object
id Int64
name string
score Float64
active boolean
dtype: object
| Goal | Code |
|---|---|
| Text to integer / float | df["col"].astype(int) / .astype(float) |
| Several columns | df.astype({"a": "int64", "b": "float64"}) |
| Messy numbers | pd.to_numeric(df["col"], errors="coerce") |
| Integers with NaN | df["col"].astype("Int64") |
| Dates | pd.to_datetime(df["col"], format="%Y-%m-%d") |
| Repeated labels | df["col"].astype("category") |
| Automatic | df.convert_dtypes() |
Related pandas tutorials:
- Convert floats to integers in pandas
- Remove all non-numeric characters in pandas
- Read a CSV file with pandas
- Convert a pandas DataFrame to JSON
Frequently asked questions
How do I change the data type of a column in pandas?
Use astype() and assign the result back: df["col"] = df["col"].astype("int64").
How do I change multiple column types at once?
Pass a dictionary: df = df.astype({"a": "int64", "b": "float64"}).
Why does astype(float) raise ValueError?
At least one value is not a number (for example "N/A" or "1,200"). Clean it, or use pd.to_numeric(..., errors="coerce").
How do I convert a column with NaN to integers?
Use the nullable type astype("Int64"), or fill the missing values first with fillna().
How do I change a column to datetime?
df["col"] = pd.to_datetime(df["col"]), with format= and errors="coerce" as needed.
When should I use the category dtype?
For columns with few distinct repeated values; it saves memory and speeds up grouping.
Bijay Kumar is a 13-time Microsoft MVP with more than 18 years in software development, and the founder of Python Guides and TSinfo Technologies. He started out building .NET and SharePoint solutions at HP, TCS and