DATA ANALYSIS

Pandas Tutorials

Pandas is how Python handles table-shaped data — spreadsheets, CSV exports, database queries. These 95 tutorials work through it by task: loading files, selecting rows, cleaning up the mess, and grouping it into an answer.

  • 95 tutorials
  • 14 topics
  • CSV to answer
Four real lines: load, group, total, sortRead the groupby tutorial

What Pandas is for

Pandas gives Python a DataFrame: a table with named columns, where each column has its own type. If you have ever built a pivot table or written a SELECT ... GROUP BY, you already know what it is for. The difference is that the whole thing is code, so it runs again next month without you clicking through it.

Nearly all Pandas work is four jobs in a row. Load the data from a CSV, Excel file, database or JSON payload. Select the rows and columns you care about. Clean it — missing values, duplicates, numbers stored as text, column names with trailing spaces. Then group and aggregate to get the number you were after.

Cleaning is the part nobody warns you about, and it is where most of the time goes. A column that looks numeric but is object dtype, a date column read as a string, a total that is wrong because of six duplicate rows — the tutorials below are grouped so you can go straight to whichever of those is biting you.

Install it and load a file

Pandas pulls in NumPy automatically, so one pip install is enough. Add openpyxl if you need to read or write .xlsx files, since Pandas does not bundle an Excel engine.

Once it is installed, read_csv() is the door in. Give it a path and it gives you a DataFrame, guessing the types as it goes — which is convenient right up until it guesses wrong.

If a column of numbers refuses to add up, check df.dtypes first. A single stray value such as "1,200" or "N/A" makes Pandas read the whole column as object, and arithmetic on it silently does the wrong thing.

Check the shape and the dtypes before anything elseFull read_csv guide

If you are starting today

A sensible order to learn Pandas in

These six steps mirror how a real analysis runs, so you can stop after any of them and still have something that works.

  1. 1

    Read a CSV into a DataFrame

    Load a file, look at its shape and check the column types before you trust anything in it.

  2. 2

    Select rows and columns

    loc, iloc and boolean filters — the one thing you will use in every single script.

  3. 3

    Filter on real conditions

    Several conditions at once, and why Python’s and does not work here.

  4. 4

    Clean the data

    Missing values, duplicate rows and numbers trapped in text columns.

  5. 5

    Group and aggregate

    Split by a column, aggregate each group, and get the summary you actually wanted.

  6. 6

    Write the result out

    Back to CSV or Excel, without the index column nobody asked for.

Quick reference

The Pandas calls you will use most

Thirteen lines that cover the bulk of everyday analysis, in roughly the order you would use them.

TaskCodeWorth knowing
Read a CSVpd.read_csv("sales.csv")index_col=0 absorbs an existing index column.
First few rowsdf.head()Pass a number for more than five.
Rows and columnsdf.shapeReturns a (rows, columns) tuple.
Column typesdf.dtypesCheck this before trusting any arithmetic.
One columndf["amount"]Gives a Series, not a DataFrame.
Filter rowsdf[df.amount > 100]Join conditions with &, never and.
Select by labeldf.loc[0, "amount"]iloc does the same by position.
Drop missing rowsdf.dropna()subset= limits it to certain columns.
Fill missing valuesdf.fillna(0)Consider the median rather than zero.
Group and totaldf.groupby("region")["amount"].sum()The single most useful line in Pandas.
Sortdf.sort_values("amount", ascending=False)Takes a list for several columns.
Rename a columndf.rename(columns={"amt": "amount"})Returns a copy unless you pass inplace=True.
Write a CSVdf.to_csv("out.csv", index=False)index=False stops the extra unnamed column.

Every tutorial, by topic

Every Pandas tutorial, grouped by what you are trying to do

These 95 tutorials used to sit in one undivided list. Same tutorials, now sorted into the fourteen jobs that make up most Pandas work.

Reading and writing files

7

Getting data in and out: CSV, Excel, JSON, text files and back again — including writing a DataFrame out without the index column.

Plotting from pandas

3

Charts straight off a DataFrame with .plot(), which calls Matplotlib for you: bar plots, scatter plots and crosstab visuals.

Creating DataFrames

3

Building a DataFrame by hand from dictionaries and lists of dictionaries — how most test data and small lookup tables start.

Selecting rows and columns

6

loc, iloc, boolean filtering and multiple conditions, plus grabbing the first row or the index of matching rows.

Missing data and cleaning

4

Filling or removing NaN values, finding and counting duplicates, and getting the unique values in a column.

Grouping and aggregation

5

groupby with aggregation functions, pivot tables, counting rows and counting unique values — where raw rows turn into an answer.

Renaming and formatting

5

Renaming columns, setting column names, showing every column instead of Pandas truncating the output, and taking the first n rows.

Index handling

5

Getting index values, setting a column as the index, resetting it, and using crosstab — the source of a surprising number of confusing bugs.

Looping and applying functions

5

iterrows, apply and lambda functions for row-by-row work — plus when a vectorised operation would be far faster.

Inspecting a DataFrame

6

Row counts, length, whether it is empty, whether a column exists, and the difference between a Series and a DataFrame.

Keep going

What to learn next to it

Pandas sits in the middle of the data stack. These are the libraries on either side of it.

Questions people ask

Frequently asked questions

Do I need to learn NumPy before Pandas?

No. Pandas is built on NumPy, but you can do real work for months without touching it directly. Pick up NumPy when you start needing element-wise maths across whole columns, or when an error message mentions ndarray and you want to know why.

What is the difference between loc and iloc?

loc selects by label — the index values and column names. iloc selects by integer position, like a list. On a DataFrame with the default 0,1,2 index they look identical, which is exactly why people get caught out later once the index is dates or IDs.

Why is my numeric column an object dtype?

Because at least one value in it is not a number. Thousands separators like "1,200", currency symbols, empty strings and "N/A" all force the column to object. Use pd.to_numeric(df["col"], errors="coerce") to convert and turn the offenders into NaN so you can see them.

How do I filter on more than one condition?

Wrap each condition in brackets and join them with & or |, not and or or: df[(df.amount > 100) & (df.region == "West")]. Python’s keywords cannot work element-wise, which is why they raise a truth-value error here.

Is Pandas too slow for large files?

It is fine into the low millions of rows on a normal laptop, as a rough guide, provided you avoid looping row by row. Read only the columns you need, set sensible dtypes, and use vectorised operations instead of iterrows. Past that, Polars or DuckDB are worth a look.

Why does my CSV export have an extra unnamed column?

Because the index was written out as a column. Pass df.to_csv("out.csv", index=False). If you are reading a file that already has one, pd.read_csv("in.csv", index_col=0) absorbs it instead of leaving an Unnamed: 0 column behind.

Load a file and ask it one question

Pick a CSV you actually care about. Read it, filter it, group it — that is a complete analysis and it takes about six lines.