scikit-learn accuracy_score: Usage and When It Misleads

The accuracy_score in sklearn returns the fraction of predictions that match the truth. Here’s the import and the whole call:

from sklearn.metrics import accuracy_score

accuracy_score(y_true, y_pred)      # 0.0 to 1.0

It takes the true labels first and the predictions second. Swapping them gives the same number, so the mistake hides, but every other metric in sklearn.metrics cares about the order.

Everything here ran on scikit-learn 1.9.1 on Python 3.12.5. The catch worth knowing about is further down: accuracy can look excellent while the model is useless.

The accuracy_score call and normalize

By default you get a fraction. Pass normalize=False and you get a count instead:

from sklearn.metrics import accuracy_score

y_true = [0, 1, 1, 0, 1, 1, 0, 1]
y_pred = [0, 1, 0, 0, 1, 1, 1, 1]

print("accuracy :", accuracy_score(y_true, y_pred))
print("correct  :", accuracy_score(y_true, y_pred, normalize=False), "out of", len(y_true))

# it is just this, counted for you
matches = sum(t == p for t, p in zip(y_true, y_pred))
print("by hand  :", matches / len(y_true))

Output:

accuracy : 0.75
correct  : 6.0 out of 8
by hand  : 0.75
Command Prompt showing scikit-learn accuracy_score returning a fraction and then a raw count with normalize set to False
Six of eight right: 0.75, or 6 with normalize=False.

That’s the entire function. It compares element by element and divides by the count, which is why the hand-written version above lands on the same number.

accuracy_score sample_weight and label types

Two arguments cover most of the rest. sample_weight lets some rows count for more, and labels can be anything as long as both lists use the same ones:

from sklearn.metrics import accuracy_score

y_true = [0, 1, 1, 0]
y_pred = [0, 1, 0, 1]          # the last two are wrong

print("unweighted:", accuracy_score(y_true, y_pred))

# make the third sample count for more than the rest
weights = [1, 1, 5, 1]
print("weighted  :", accuracy_score(y_true, y_pred, sample_weight=weights))

# strings and any other labels work as long as both sides agree
animals_true = ["cat", "dog", "cat", "bird"]
animals_pred = ["cat", "dog", "dog", "bird"]
print("labels    :", accuracy_score(animals_true, animals_pred))

Output:

unweighted: 0.5
weighted  : 0.25
labels    : 0.75

Strings, integers, booleans all work. What doesn’t work is a mismatch, such as predicting "1" against a true label of 1, which counts as wrong without complaint.

When accuracy lies to you

This is the part the metric’s reputation rests on, and it’s easy to demonstrate.

Take 1,000 transactions where 20 are fraud. Build a model that simply never predicts fraud:

import numpy as np
from sklearn.metrics import accuracy_score, balanced_accuracy_score, confusion_matrix

# 1000 transactions, 20 of them fraudulent
y_true = np.zeros(1000, dtype=int)
y_true[:20] = 1

# a "model" that never predicts fraud at all
lazy = np.zeros(1000, dtype=int)

print("accuracy          :", accuracy_score(y_true, lazy))
print("balanced accuracy :", balanced_accuracy_score(y_true, lazy))
print("fraud caught      :", int(((y_true == 1) & (lazy == 1)).sum()), "of", int((y_true == 1).sum()))
print("confusion matrix  :")
print(confusion_matrix(y_true, lazy))

Output:

accuracy          : 0.98
balanced accuracy : 0.5
fraud caught      : 0 of 20
confusion matrix  :
[[980   0]
 [ 20   0]]
Command Prompt showing a model with 98 percent accuracy that catches zero fraud cases, with the balanced accuracy at 0.5 and the confusion matrix
98% accurate, and it caught none of the 20 frauds.

The model has learned nothing and would pass a 95% threshold in a review. Balanced accuracy tells the truth at 0.5, which is what you’d get from a coin toss.

The confusion matrix is the giveaway: a whole column of zeros means a class was never predicted at all.

What should you use instead of accuracy?

Accuracy isn’t wrong, it’s just incomplete. classification_report gives you precision, recall and F1 per class in one call:

import numpy as np
from sklearn.metrics import accuracy_score, classification_report

y_true = np.zeros(1000, dtype=int)
y_true[:20] = 1
lazy = np.zeros(1000, dtype=int)

print("accuracy:", accuracy_score(y_true, lazy))
print(classification_report(y_true, lazy, target_names=["legit", "fraud"], zero_division=0))

Output:

accuracy: 0.98
              precision    recall  f1-score   support

       legit       0.98      1.00      0.99       980
       fraud       0.00      0.00      0.00        20

    accuracy                           0.98      1000
   macro avg       0.49      0.50      0.49      1000
weighted avg       0.96      0.98      0.97      1000
Command Prompt showing a scikit-learn classification report with zero precision and recall for the fraud class next to a high accuracy figure
Recall 0.00 on the class that matters. Accuracy never showed that.
MetricAnswersReach for it when
accuracy_scoreWhat fraction did I get right?Classes are roughly balanced
balanced_accuracy_scoreAverage recall across classesOne class is rare
precision_scoreOf my positives, how many were real?False alarms are costly
recall_scoreOf the real positives, how many did I catch?Misses are costly
f1_scoreThe balance of those twoYou need one number

Pick the metric from the cost of being wrong. Missing a fraud and flagging a good customer are not the same mistake, and accuracy treats them as if they were.

accuracy_score or model.score()?

Every scikit-learn classifier has a .score() method, and for classifiers it is accuracy:

from sklearn.datasets import load_iris
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
from sklearn.model_selection import train_test_split

X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)

model = LogisticRegression(max_iter=200).fit(X_train, y_train)
predictions = model.predict(X_test)

print("accuracy_score :", accuracy_score(y_test, predictions))
print("model.score()  :", model.score(X_test, y_test))
print("the same call  :", accuracy_score(y_test, predictions) == model.score(X_test, y_test))

Output:

accuracy_score : 1.0
model.score()  : 1.0
the same call  : True

Use .score() for a quick look, and accuracy_score when you already have predictions in hand or want to pass sample_weight.

Train vs test accuracy and cross-validation

A single accuracy figure hides two things: whether the model memorised the training set, and how much the number moves when you resample:

from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
from sklearn.model_selection import cross_val_score, train_test_split
from sklearn.tree import DecisionTreeClassifier

# a realistic dataset: some noise, some uninformative columns
X, y = make_classification(n_samples=400, n_features=12, n_informative=4,
                           flip_y=0.15, random_state=0)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)

tree = DecisionTreeClassifier(random_state=0).fit(X_train, y_train)
print("train accuracy:", accuracy_score(y_train, tree.predict(X_train)))
print("test  accuracy:", accuracy_score(y_test, tree.predict(X_test)))

folds = cross_val_score(LogisticRegression(max_iter=500), X, y, cv=5, scoring="accuracy")
print("5-fold scores :", folds.round(3).tolist())
print("mean          :", folds.mean().round(4), "+/-", folds.std().round(4))

Output:

train accuracy: 1.0
test  accuracy: 0.675
5-fold scores : [0.612, 0.638, 0.7, 0.7, 0.725]
mean          : 0.675 +/- 0.0426
Command Prompt showing a decision tree scoring perfect accuracy on training data but lower on test data, next to five cross-validation fold scores
A perfect training score and a lower test score, plus what five different splits actually give.

The tree scores a flawless 1.0 on data it has already seen and 0.675 on data it hasn’t. Only the second number means anything.

The fold scores matter just as much. These five range from 0.612 to 0.725, so any single train/test split was partly luck.

Quote the mean with its spread. One accuracy figure with no error bar invites the reader to believe a precision you don’t have.

accuracy_score on multilabel data

With multilabel data, accuracy_score means subset accuracy: a row counts only if every single label is right.

import numpy as np
from sklearn.metrics import accuracy_score

# three samples, three possible tags each
y_true = np.array([[1, 0, 1],
                   [0, 1, 1],
                   [1, 1, 0]])

y_pred = np.array([[1, 0, 1],      # exactly right
                   [0, 1, 0],      # one tag missed
                   [1, 1, 1]])     # one tag too many

print("accuracy_score :", accuracy_score(y_true, y_pred), "<- whole rows only")
print("per-label match:", (y_true == y_pred).mean(), "<- what people usually expect")

Output:

accuracy_score : 0.3333333333333333 <- whole rows only
per-label match: 0.7777777777777778 <- what people usually expect
Command Prompt showing scikit-learn accuracy_score returning 0.33 for multilabel data while the per-label match rate is 0.78
One row of three was perfect, so subset accuracy is 0.33 even though most labels were right.

If you wanted the per-label figure, compute it yourself or use hamming_loss. This surprises people far more than it should, rather like NumPy quietly promoting a dtype.

More Python data guides worth a look:

Frequently asked questions

How do I import accuracy_score?

from sklearn.metrics import accuracy_score. It lives in sklearn.metrics alongside the other scoring functions, which are listed in the accuracy_score reference.

What does accuracy_score return?

A float between 0 and 1, the fraction of predictions that matched. Pass normalize=False to get the raw count of correct predictions instead.

What is the argument order?

accuracy_score(y_true, y_pred), truth first. The result is the same either way round, but precision, recall and the confusion matrix are not.

Why is my accuracy high but the model useless?

Your classes are imbalanced. Predicting the majority class every time scores well; check balanced_accuracy_score and the confusion matrix.

Is accuracy_score the same as model.score()?

For classifiers, yes. .score() predicts and then computes accuracy. Regressors return R-squared from .score() instead.

How does accuracy_score handle multilabel data?

As subset accuracy: a sample counts only when every label matches. Use hamming_loss if you want credit for partly correct rows.

Can I weight some samples more heavily?

Yes, pass sample_weight with one weight per sample. The result becomes the weighted fraction correct.