How to Plot a Line of Best Fit in Python (Matplotlib)

To draw a line of best fit in Python with Matplotlib, calculate the slope and intercept with np.polyfit(x, y, 1), then plot slope * x + intercept on top of a scatter plot of the data. This matplotlib best fit line tutorial shows the full code, how to print the equation and R², how to get statistics with scipy.stats.linregress, how to draw a line from a known slope and intercept, curves of best fit, seaborn’s regplot, and what to do about missing values.

Tested with Python 3.12.5, Matplotlib 3.11.2, NumPy 2.5.3, SciPy 1.18.1 and seaborn 0.13.2. Charts are real Matplotlib windows and the numbers come from real runs in the Windows Command Prompt. References: numpy.polyfit, scipy.stats.linregress and Axes.axline.

Scatter plot with a line of best fit

A line of best fit (linear regression line) is the straight line that keeps the squared vertical distances to all points as small as possible. np.polyfit with degree 1 returns its slope and intercept:

import matplotlib.pyplot as plt
import numpy as np

# advertising spend (thousand USD) and sales (thousand units) for 12 months
ad_spend = np.array([2.1, 3.4, 4.0, 4.8, 5.5, 6.1, 6.9, 7.4, 8.2, 9.0, 9.6, 10.5])
sales = np.array([14.0, 17.9, 18.2, 22.5, 23.1, 26.8, 27.0, 30.9, 31.2, 35.8, 35.1, 39.6])

slope, intercept = np.polyfit(ad_spend, sales, 1)     # degree 1 = straight line

plt.scatter(ad_spend, sales, label="Monthly data")
plt.plot(ad_spend, slope * ad_spend + intercept, color="red", label="Line of best fit")
plt.xlabel("Advertising spend (thousand USD)")
plt.ylabel("Sales (thousand units)")
plt.legend()
plt.show()
Matplotlib scatter plot of advertising spend against sales with a red line of best fit calculated with numpy polyfit
Twelve months of data and the line of best fit from np.polyfit.

Print the slope, intercept, equation and R²

The slope says how much y changes when x grows by 1; R² (0 to 1) says how much of the variation the line explains:

import numpy as np

# advertising spend (thousand USD) and sales (thousand units) for 12 months
ad_spend = np.array([2.1, 3.4, 4.0, 4.8, 5.5, 6.1, 6.9, 7.4, 8.2, 9.0, 9.6, 10.5])
sales = np.array([14.0, 17.9, 18.2, 22.5, 23.1, 26.8, 27.0, 30.9, 31.2, 35.8, 35.1, 39.6])

slope, intercept = np.polyfit(ad_spend, sales, 1)
predicted = slope * ad_spend + intercept
r_squared = 1 - np.sum((sales - predicted) ** 2) / np.sum((sales - sales.mean()) ** 2)

print(f"slope     = {slope:.3f}")
print(f"intercept = {intercept:.3f}")
print(f"equation  : y = {slope:.2f}x + {intercept:.2f}")
print(f"R squared = {r_squared:.4f}")
print(f"forecast for 12k spend: {slope * 12 + intercept:.1f} thousand units")

Output:

slope     = 3.030
intercept = 7.276
equation  : y = 3.03x + 7.28
R squared = 0.9842
forecast for 12k spend: 43.6 thousand units
Command Prompt output of numpy polyfit showing the slope, intercept, line equation, R squared and a sales forecast
Slope, intercept, equation and R² printed in the Command Prompt.

Line of best fit with scipy.stats.linregress

linregress returns the same line plus the correlation coefficient r, the p-value and standard errors in one call:

import numpy as np
from scipy import stats

# advertising spend (thousand USD) and sales (thousand units) for 12 months
ad_spend = np.array([2.1, 3.4, 4.0, 4.8, 5.5, 6.1, 6.9, 7.4, 8.2, 9.0, 9.6, 10.5])
sales = np.array([14.0, 17.9, 18.2, 22.5, 23.1, 26.8, 27.0, 30.9, 31.2, 35.8, 35.1, 39.6])

result = stats.linregress(ad_spend, sales)

print(f"slope     = {result.slope:.3f}")
print(f"intercept = {result.intercept:.3f}")
print(f"r         = {result.rvalue:.4f}  (R squared = {result.rvalue ** 2:.4f})")
print(f"p-value   = {result.pvalue:.2e}")
print(f"std error = {result.stderr:.3f}")

Output:

slope     = 3.030
intercept = 7.276
r         = 0.9921  (R squared = 0.9842)
p-value   = 2.43e-10
std error = 0.121
Command Prompt output of scipy.stats.linregress with slope, intercept, r value, R squared, p-value and standard error
The same slope and intercept as polyfit, plus statistics.

Show the equation and R² on the plot

import matplotlib.pyplot as plt
import numpy as np
from scipy import stats

# advertising spend (thousand USD) and sales (thousand units) for 12 months
ad_spend = np.array([2.1, 3.4, 4.0, 4.8, 5.5, 6.1, 6.9, 7.4, 8.2, 9.0, 9.6, 10.5])
sales = np.array([14.0, 17.9, 18.2, 22.5, 23.1, 26.8, 27.0, 30.9, 31.2, 35.8, 35.1, 39.6])

res = stats.linregress(ad_spend, sales)
x_line = np.linspace(0, 11, 100)

fig, ax = plt.subplots(figsize=(7, 4.5))
ax.scatter(ad_spend, sales, color="tab:blue", zorder=3)
ax.plot(x_line, res.slope * x_line + res.intercept, color="tab:red",
        label=f"y = {res.slope:.2f}x + {res.intercept:.2f}   (R² = {res.rvalue ** 2:.3f})")
ax.set_xlabel("Advertising spend (thousand USD)")
ax.set_ylabel("Sales (thousand units)")
ax.legend(loc="upper left")
ax.grid(alpha=0.3)
plt.tight_layout()
plt.show()
Matplotlib scatter plot with a red regression line whose equation and R squared value are shown in the legend
The equation and R² in the legend; the line is extended over the whole x-range with np.linspace.

Plot a line from a slope and intercept

If you already know the slope and intercept, you do not need x-values: ax.axline() (Matplotlib 3.3+) draws an infinite line through a point with a given slope, and it always reaches the edges of the plot:

import matplotlib.pyplot as plt

slope, intercept = 1.5, 2

fig, ax = plt.subplots(figsize=(6, 4.5))
ax.axline((0, intercept), slope=slope, color="purple", label="y = 1.5x + 2")   # infinite line
ax.axline((0, 0), slope=-0.5, color="gray", linestyle="--", label="y = -0.5x")
ax.set_xlim(-2, 6)
ax.set_ylim(-4, 12)
ax.axhline(0, color="black", linewidth=0.6)
ax.axvline(0, color="black", linewidth=0.6)
ax.legend()
plt.show()
Matplotlib plot of two infinite lines drawn with axline from a slope and intercept, y = 1.5x + 2 and y = -0.5x
ax.axline((0, intercept), slope=slope).

Curve of best fit (polynomial)

When the points bend, fit a polynomial of higher degree. np.poly1d turns the coefficients into a function you can call:

import matplotlib.pyplot as plt
import numpy as np

hours = np.array([0, 1, 2, 3, 4, 5, 6, 7, 8])
temperature = np.array([12.0, 14.8, 18.1, 20.2, 21.9, 22.4, 21.8, 20.1, 17.6])

coeffs = np.polyfit(hours, temperature, 2)        # degree 2 = parabola
curve = np.poly1d(coeffs)                          # callable polynomial
print("coefficients (a, b, c):", np.round(coeffs, 3))   # y = a*x**2 + b*x + c
print("temperature at 4.5 h:", round(curve(4.5), 1))

x = np.linspace(0, 8, 200)
plt.scatter(hours, temperature, label="Measured")
plt.plot(x, curve(x), color="darkorange", label="Curve of best fit (degree 2)")
plt.xlabel("Hours after 6 am")
plt.ylabel("Temperature (°C)")
plt.legend()
plt.show()

Output:

coefficients (a, b, c): [-0.442  4.333 11.449]
temperature at 4.5 h: 22.0
Matplotlib scatter plot of temperatures during the day with an orange degree 2 polynomial curve of best fit
A degree-2 curve of best fit. Keep the degree low: high degrees chase noise.

Best fit line with a confidence band in seaborn

seaborn’s regplot fits and draws the line in one call and shades a 95% confidence interval around it:

import matplotlib.pyplot as plt
import numpy as np
import seaborn as sns

# advertising spend (thousand USD) and sales (thousand units) for 12 months
ad_spend = np.array([2.1, 3.4, 4.0, 4.8, 5.5, 6.1, 6.9, 7.4, 8.2, 9.0, 9.6, 10.5])
sales = np.array([14.0, 17.9, 18.2, 22.5, 23.1, 26.8, 27.0, 30.9, 31.2, 35.8, 35.1, 39.6])

ax = sns.regplot(x=ad_spend, y=sales, ci=95, line_kws={"color": "red"})   # line + 95% confidence band
ax.set_xlabel("Advertising spend (thousand USD)")
ax.set_ylabel("Sales (thousand units)")
plt.show()
seaborn regplot scatter plot with a red regression line and a shaded 95 percent confidence interval band
sns.regplot() with its confidence band.

Missing values: np.polyfit returns nan

If x or y contains NaN, np.polyfit does not raise an error. It returns nan for the slope and intercept, and the line simply does not appear on the chart. Drop incomplete pairs first:

import numpy as np

x = np.array([1.0, 2.0, 3.0, 4.0, 5.0, 6.0])
y = np.array([2.1, 3.9, np.nan, 8.2, 9.8, 12.1])   # one missing value

print("with NaN:   ", np.polyfit(x, y, 1))          # no error, just nan

ok = ~np.isnan(x) & ~np.isnan(y)                   # keep only complete pairs
print("NaN removed:", np.polyfit(x[ok], y[ok], 1))

Output:

with NaN:    [nan nan]
NaN removed: [1.99651163 0.03255814]
Command Prompt showing numpy polyfit returning nan nan when the data contains NaN and a valid slope and intercept after removing it
With a NaN the fit is [nan nan]; after masking it works.

With pandas, df.dropna(subset=["x", "y"]) does the same.

Which method should you use?

NeedUse
Just the line on a chartnp.polyfit(x, y, 1)
r, p-value, standard errorsscipy.stats.linregress(x, y)
A line from a known slope and interceptax.axline((0, b), slope=m)
A curved trendnp.polyfit(x, y, 2) + np.poly1d
A quick chart with a confidence bandseaborn.regplot()

Related Matplotlib and NumPy tutorials:

Frequently asked questions

How do I plot a line of best fit in Matplotlib?

Get the slope and intercept with m, b = np.polyfit(x, y, 1), then call plt.scatter(x, y) and plt.plot(x, m * x + b).

How do I find the equation of the line of best fit in Python?

np.polyfit(x, y, 1) returns [slope, intercept], so the equation is y = slope * x + intercept. scipy.stats.linregress returns the same values as .slope and .intercept.

How do I plot a line with a given slope and intercept in Matplotlib?

Use ax.axline((0, intercept), slope=slope). It draws an infinite line, so you do not need to choose x-values.

How do I calculate R squared for a best fit line?

Use scipy.stats.linregress(x, y).rvalue ** 2, or compute 1 - SS_res / SS_tot from the predicted values.

Why does my line of best fit not show up?

Usually the data contains NaN: np.polyfit then returns nan without an error. Remove missing values with a mask or dropna().

How do I fit a curve instead of a straight line?

Use a higher degree: np.poly1d(np.polyfit(x, y, 2)) gives a quadratic curve you can evaluate on a np.linspace range.