The accuracy_score in sklearn returns the fraction of predictions that match the truth. Here’s the import and the whole call:
from sklearn.metrics import accuracy_score
accuracy_score(y_true, y_pred) # 0.0 to 1.0
It takes the true labels first and the predictions second. Swapping them gives the same number, so the mistake hides, but every other metric in sklearn.metrics cares about the order.
Everything here ran on scikit-learn 1.9.1 on Python 3.12.5. The catch worth knowing about is further down: accuracy can look excellent while the model is useless.
The accuracy_score call and normalize
By default you get a fraction. Pass normalize=False and you get a count instead:
from sklearn.metrics import accuracy_score
y_true = [0, 1, 1, 0, 1, 1, 0, 1]
y_pred = [0, 1, 0, 0, 1, 1, 1, 1]
print("accuracy :", accuracy_score(y_true, y_pred))
print("correct :", accuracy_score(y_true, y_pred, normalize=False), "out of", len(y_true))
# it is just this, counted for you
matches = sum(t == p for t, p in zip(y_true, y_pred))
print("by hand :", matches / len(y_true))
Output:
accuracy : 0.75
correct : 6.0 out of 8
by hand : 0.75
0.75, or 6 with normalize=False.That’s the entire function. It compares element by element and divides by the count, which is why the hand-written version above lands on the same number.
accuracy_score sample_weight and label types
Two arguments cover most of the rest. sample_weight lets some rows count for more, and labels can be anything as long as both lists use the same ones:
from sklearn.metrics import accuracy_score
y_true = [0, 1, 1, 0]
y_pred = [0, 1, 0, 1] # the last two are wrong
print("unweighted:", accuracy_score(y_true, y_pred))
# make the third sample count for more than the rest
weights = [1, 1, 5, 1]
print("weighted :", accuracy_score(y_true, y_pred, sample_weight=weights))
# strings and any other labels work as long as both sides agree
animals_true = ["cat", "dog", "cat", "bird"]
animals_pred = ["cat", "dog", "dog", "bird"]
print("labels :", accuracy_score(animals_true, animals_pred))
Output:
unweighted: 0.5
weighted : 0.25
labels : 0.75
Strings, integers, booleans all work. What doesn’t work is a mismatch, such as predicting "1" against a true label of 1, which counts as wrong without complaint.
When accuracy lies to you
This is the part the metric’s reputation rests on, and it’s easy to demonstrate.
Take 1,000 transactions where 20 are fraud. Build a model that simply never predicts fraud:
import numpy as np
from sklearn.metrics import accuracy_score, balanced_accuracy_score, confusion_matrix
# 1000 transactions, 20 of them fraudulent
y_true = np.zeros(1000, dtype=int)
y_true[:20] = 1
# a "model" that never predicts fraud at all
lazy = np.zeros(1000, dtype=int)
print("accuracy :", accuracy_score(y_true, lazy))
print("balanced accuracy :", balanced_accuracy_score(y_true, lazy))
print("fraud caught :", int(((y_true == 1) & (lazy == 1)).sum()), "of", int((y_true == 1).sum()))
print("confusion matrix :")
print(confusion_matrix(y_true, lazy))
Output:
accuracy : 0.98
balanced accuracy : 0.5
fraud caught : 0 of 20
confusion matrix :
[[980 0]
[ 20 0]]
The model has learned nothing and would pass a 95% threshold in a review. Balanced accuracy tells the truth at 0.5, which is what you’d get from a coin toss.
The confusion matrix is the giveaway: a whole column of zeros means a class was never predicted at all.
What should you use instead of accuracy?
Accuracy isn’t wrong, it’s just incomplete. classification_report gives you precision, recall and F1 per class in one call:
import numpy as np
from sklearn.metrics import accuracy_score, classification_report
y_true = np.zeros(1000, dtype=int)
y_true[:20] = 1
lazy = np.zeros(1000, dtype=int)
print("accuracy:", accuracy_score(y_true, lazy))
print(classification_report(y_true, lazy, target_names=["legit", "fraud"], zero_division=0))
Output:
accuracy: 0.98
precision recall f1-score support
legit 0.98 1.00 0.99 980
fraud 0.00 0.00 0.00 20
accuracy 0.98 1000
macro avg 0.49 0.50 0.49 1000
weighted avg 0.96 0.98 0.97 1000
0.00 on the class that matters. Accuracy never showed that.| Metric | Answers | Reach for it when |
|---|---|---|
accuracy_score | What fraction did I get right? | Classes are roughly balanced |
balanced_accuracy_score | Average recall across classes | One class is rare |
precision_score | Of my positives, how many were real? | False alarms are costly |
recall_score | Of the real positives, how many did I catch? | Misses are costly |
f1_score | The balance of those two | You need one number |
Pick the metric from the cost of being wrong. Missing a fraud and flagging a good customer are not the same mistake, and accuracy treats them as if they were.
accuracy_score or model.score()?
Every scikit-learn classifier has a .score() method, and for classifiers it is accuracy:
from sklearn.datasets import load_iris
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
from sklearn.model_selection import train_test_split
X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
model = LogisticRegression(max_iter=200).fit(X_train, y_train)
predictions = model.predict(X_test)
print("accuracy_score :", accuracy_score(y_test, predictions))
print("model.score() :", model.score(X_test, y_test))
print("the same call :", accuracy_score(y_test, predictions) == model.score(X_test, y_test))
Output:
accuracy_score : 1.0
model.score() : 1.0
the same call : True
Use .score() for a quick look, and accuracy_score when you already have predictions in hand or want to pass sample_weight.
Train vs test accuracy and cross-validation
A single accuracy figure hides two things: whether the model memorised the training set, and how much the number moves when you resample:
from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
from sklearn.model_selection import cross_val_score, train_test_split
from sklearn.tree import DecisionTreeClassifier
# a realistic dataset: some noise, some uninformative columns
X, y = make_classification(n_samples=400, n_features=12, n_informative=4,
flip_y=0.15, random_state=0)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
tree = DecisionTreeClassifier(random_state=0).fit(X_train, y_train)
print("train accuracy:", accuracy_score(y_train, tree.predict(X_train)))
print("test accuracy:", accuracy_score(y_test, tree.predict(X_test)))
folds = cross_val_score(LogisticRegression(max_iter=500), X, y, cv=5, scoring="accuracy")
print("5-fold scores :", folds.round(3).tolist())
print("mean :", folds.mean().round(4), "+/-", folds.std().round(4))
Output:
train accuracy: 1.0
test accuracy: 0.675
5-fold scores : [0.612, 0.638, 0.7, 0.7, 0.725]
mean : 0.675 +/- 0.0426
The tree scores a flawless 1.0 on data it has already seen and 0.675 on data it hasn’t. Only the second number means anything.
The fold scores matter just as much. These five range from 0.612 to 0.725, so any single train/test split was partly luck.
Quote the mean with its spread. One accuracy figure with no error bar invites the reader to believe a precision you don’t have.
accuracy_score on multilabel data
With multilabel data, accuracy_score means subset accuracy: a row counts only if every single label is right.
import numpy as np
from sklearn.metrics import accuracy_score
# three samples, three possible tags each
y_true = np.array([[1, 0, 1],
[0, 1, 1],
[1, 1, 0]])
y_pred = np.array([[1, 0, 1], # exactly right
[0, 1, 0], # one tag missed
[1, 1, 1]]) # one tag too many
print("accuracy_score :", accuracy_score(y_true, y_pred), "<- whole rows only")
print("per-label match:", (y_true == y_pred).mean(), "<- what people usually expect")
Output:
accuracy_score : 0.3333333333333333 <- whole rows only
per-label match: 0.7777777777777778 <- what people usually expect
If you wanted the per-label figure, compute it yourself or use hamming_loss. This surprises people far more than it should, rather like NumPy quietly promoting a dtype.
More Python data guides worth a look:
- NumPy data types
- Create a 2D array with NumPy
- Find the maximum value in an array
- SciPy root finding
- Round numbers to 2 decimal places
Frequently asked questions
How do I import accuracy_score?
from sklearn.metrics import accuracy_score. It lives in sklearn.metrics alongside the other scoring functions, which are listed in the accuracy_score reference.
What does accuracy_score return?
A float between 0 and 1, the fraction of predictions that matched. Pass normalize=False to get the raw count of correct predictions instead.
What is the argument order?
accuracy_score(y_true, y_pred), truth first. The result is the same either way round, but precision, recall and the confusion matrix are not.
Why is my accuracy high but the model useless?
Your classes are imbalanced. Predicting the majority class every time scores well; check balanced_accuracy_score and the confusion matrix.
Is accuracy_score the same as model.score()?
For classifiers, yes. .score() predicts and then computes accuracy. Regressors return R-squared from .score() instead.
How does accuracy_score handle multilabel data?
As subset accuracy: a sample counts only when every label matches. Use hamming_loss if you want credit for partly correct rows.
Can I weight some samples more heavily?
Yes, pass sample_weight with one weight per sample. The result becomes the weighted fraction correct.
Bijay Kumar is a 13-time Microsoft MVP with more than 18 years in software development, and the founder of Python Guides and TSinfo Technologies. He started out building .NET and SharePoint solutions at HP, TCS and KPIT before moving into Python, machine learning and AI, and he also builds web apps with TypeScript and React. He writes the tutorials here himself, and every example is run before publishing so you see the real output. More about Bijay · Microsoft MVP profile · LinkedIn