DATA ANALYSIS
Pandas Tutorials
Pandas is how Python handles table-shaped data — spreadsheets, CSV exports, database queries. These 95 tutorials work through it by task: loading files, selecting rows, cleaning up the mess, and grouping it into an answer.
- 95 tutorials
- 14 topics
- CSV to answer
What Pandas is for
Pandas gives Python a DataFrame: a table with named columns, where each column has its own type. If you have ever built a pivot table or written a SELECT ... GROUP BY, you already know what it is for. The difference is that the whole thing is code, so it runs again next month without you clicking through it.
Nearly all Pandas work is four jobs in a row. Load the data from a CSV, Excel file, database or JSON payload. Select the rows and columns you care about. Clean it — missing values, duplicates, numbers stored as text, column names with trailing spaces. Then group and aggregate to get the number you were after.
Cleaning is the part nobody warns you about, and it is where most of the time goes. A column that looks numeric but is object dtype, a date column read as a string, a total that is wrong because of six duplicate rows — the tutorials below are grouped so you can go straight to whichever of those is biting you.
Install it and load a file
Pandas pulls in NumPy automatically, so one pip install is enough. Add openpyxl if you need to read or write .xlsx files, since Pandas does not bundle an Excel engine.
Once it is installed, read_csv() is the door in. Give it a path and it gives you a DataFrame, guessing the types as it goes — which is convenient right up until it guesses wrong.
If a column of numbers refuses to add up, check df.dtypes first. A single stray value such as "1,200" or "N/A" makes Pandas read the whole column as object, and arithmetic on it silently does the wrong thing.
If you are starting today
A sensible order to learn Pandas in
These six steps mirror how a real analysis runs, so you can stop after any of them and still have something that works.
- 1
Read a CSV into a DataFrame
Load a file, look at its shape and check the column types before you trust anything in it.
- 2
Select rows and columns
loc,ilocand boolean filters — the one thing you will use in every single script. - 3
Filter on real conditions
Several conditions at once, and why Python’s
anddoes not work here. - 4
Clean the data
Missing values, duplicate rows and numbers trapped in text columns.
- 5
Group and aggregate
Split by a column, aggregate each group, and get the summary you actually wanted.
- 6
Write the result out
Back to CSV or Excel, without the index column nobody asked for.
Quick reference
The Pandas calls you will use most
Thirteen lines that cover the bulk of everyday analysis, in roughly the order you would use them.
| Task | Code | Worth knowing |
|---|---|---|
| Read a CSV | pd.read_csv("sales.csv") | index_col=0 absorbs an existing index column. |
| First few rows | df.head() | Pass a number for more than five. |
| Rows and columns | df.shape | Returns a (rows, columns) tuple. |
| Column types | df.dtypes | Check this before trusting any arithmetic. |
| One column | df["amount"] | Gives a Series, not a DataFrame. |
| Filter rows | df[df.amount > 100] | Join conditions with &, never and. |
| Select by label | df.loc[0, "amount"] | iloc does the same by position. |
| Drop missing rows | df.dropna() | subset= limits it to certain columns. |
| Fill missing values | df.fillna(0) | Consider the median rather than zero. |
| Group and total | df.groupby("region")["amount"].sum() | The single most useful line in Pandas. |
| Sort | df.sort_values("amount", ascending=False) | Takes a list for several columns. |
| Rename a column | df.rename(columns={"amt": "amount"}) | Returns a copy unless you pass inplace=True. |
| Write a CSV | df.to_csv("out.csv", index=False) | index=False stops the extra unnamed column. |
Every tutorial, by topic
Every Pandas tutorial, grouped by what you are trying to do
These 95 tutorials used to sit in one undivided list. Same tutorials, now sorted into the fourteen jobs that make up most Pandas work.
Reading and writing files
7Getting data in and out: CSV, Excel, JSON, text files and back again — including writing a DataFrame out without the index column.
Plotting from pandas
3Charts straight off a DataFrame with .plot(), which calls Matplotlib for you: bar plots, scatter plots and crosstab visuals.
Creating DataFrames
3Building a DataFrame by hand from dictionaries and lists of dictionaries — how most test data and small lookup tables start.
Selecting rows and columns
6loc, iloc, boolean filtering and multiple conditions, plus grabbing the first row or the index of matching rows.
Adding and removing
16Adding rows and columns, dropping rows, duplicates, unnamed columns, header rows and non-numeric data you do not want.
Missing data and cleaning
4Filling or removing NaN values, finding and counting duplicates, and getting the unique values in a column.
Grouping and aggregation
5groupby with aggregation functions, pivot tables, counting rows and counting unique values — where raw rows turn into an answer.
Types and conversion
16Converting between DataFrames, dictionaries, lists and NumPy arrays, floats to integers, Series to DataFrame, and on into a TensorFlow dataset.
Renaming and formatting
5Renaming columns, setting column names, showing every column instead of Pandas truncating the output, and taking the first n rows.
Index handling
5Getting index values, setting a column as the index, resetting it, and using crosstab — the source of a surprising number of confusing bugs.
Looping and applying functions
5iterrows, apply and lambda functions for row-by-row work — plus when a vectorised operation would be far faster.
Replacing and updating values
4Replacing one value or many, replacing conditionally, and updating a column in place.
Inspecting a DataFrame
6Row counts, length, whether it is empty, whether a column exists, and the difference between a Series and a DataFrame.
More Pandas tutorials
10Sorting, merging, dates as an index, splitting a column and comparing two DataFrames, plus interview questions.
Keep going
What to learn next to it
Pandas sits in the middle of the data stack. These are the libraries on either side of it.
Questions people ask
Frequently asked questions
Do I need to learn NumPy before Pandas?
No. Pandas is built on NumPy, but you can do real work for months without touching it directly. Pick up NumPy when you start needing element-wise maths across whole columns, or when an error message mentions ndarray and you want to know why.
What is the difference between loc and iloc?
loc selects by label — the index values and column names. iloc selects by integer position, like a list. On a DataFrame with the default 0,1,2 index they look identical, which is exactly why people get caught out later once the index is dates or IDs.
Why is my numeric column an object dtype?
Because at least one value in it is not a number. Thousands separators like "1,200", currency symbols, empty strings and "N/A" all force the column to object. Use pd.to_numeric(df["col"], errors="coerce") to convert and turn the offenders into NaN so you can see them.
How do I filter on more than one condition?
Wrap each condition in brackets and join them with & or |, not and or or: df[(df.amount > 100) & (df.region == "West")]. Python’s keywords cannot work element-wise, which is why they raise a truth-value error here.
Is Pandas too slow for large files?
It is fine into the low millions of rows on a normal laptop, as a rough guide, provided you avoid looping row by row. Read only the columns you need, set sensible dtypes, and use vectorised operations instead of iterrows. Past that, Polars or DuckDB are worth a look.
Why does my CSV export have an extra unnamed column?
Because the index was written out as a column. Pass df.to_csv("out.csv", index=False). If you are reading a file that already has one, pd.read_csv("in.csv", index_col=0) absorbs it instead of leaving an Unnamed: 0 column behind.
Load a file and ask it one question
Pick a CSV you actually care about. Read it, filter it, group it — that is a complete analysis and it takes about six lines.
