CLASSICAL MACHINE LEARNING

Scikit-learn Tutorials

Scikit-learn is the right tool for most machine learning problems that are not deep learning — classification, regression and clustering on table-shaped data, with one consistent interface across every model.

  • One interface
  • No GPU needed
  • Minutes to a model
fit, predict, score — the same for every modelRead the tutorial

What scikit-learn is for

Scikit-learn covers classical machine learning: the algorithms that work on rows and columns rather than images or text. Logistic regression, decision trees, random forests, gradient boosting, k-means, support vector machines — along with the unglamorous but essential work of scaling features, encoding categories and splitting data honestly.

Its best feature is consistency. Every model exposes the same three methods: fit() to learn, predict() to apply, and score() to check. Swapping a logistic regression for a random forest is a one-line change, which makes comparing several approaches genuinely quick.

For most business problems on tabular data it is also the better choice, not just the easier one. A gradient-boosted tree will usually beat a neural network on a spreadsheet-shaped dataset, trains in seconds on a laptop, and produces a model you can explain to the person who has to act on it.

The honest caveat is that scikit-learn is not for deep learning. Images, audio and language belong in Keras or PyTorch. The dividing line is roughly whether your data has meaningful columns — if it does, start here.

Install it and fit a model

Scikit-learn installs with pip and brings NumPy and SciPy with it. The package is named scikit-learn but imported as sklearn, which catches people out the first time.

Split your data before you do anything else. Fitting on everything and reporting the score on the same rows produces a number that looks excellent and means nothing.

Scale your features before any distance-based or regularised model — k-nearest neighbours, SVMs, logistic regression. Fit the scaler on the training set only, then apply it to the test set, or information leaks across the split and your score is optimistic.

Install scikit-learn, import sklearnFit your first model

If you are starting today

A sensible order to learn scikit-learn in

Five steps from a first model to reading its results honestly. Accuracy comes late on purpose — it is the number people trust too early.

  1. 1

    Fit your first classifier

    Logistic regression on a split dataset: the fit, predict and score pattern every other model follows.

  2. 2

    Understand regression

    Predicting a number rather than a category, and when each is the right framing.

  3. 3

    Check accuracy properly

    What the score actually measures, and why it misleads on imbalanced data.

  4. 4

    Read a confusion matrix

    Where the model is wrong matters more than how often — false positives and false negatives rarely cost the same.

  5. 5

    Prepare better features

    Most real gains come from the features, not from swapping the algorithm.

Quick reference

The scikit-learn calls you will use most

Ten lines that cover a complete modelling run, from split to score.

TaskCodeWorth knowing
Split the datatrain_test_split(X, y, test_size=0.2)Add stratify=y for imbalanced classes.
Scale featuresStandardScaler().fit_transform(X_train)Fit on train only, then transform test.
Encode categoriesOneHotEncoder(handle_unknown="ignore")Handles unseen values at predict time.
Train a modelmodel.fit(X_train, y_train)The same call for every estimator.
Predictmodel.predict(X_test)predict_proba() gives probabilities.
Quick scoremodel.score(X_test, y_test)Accuracy for classifiers, R² for regressors.
Confusion matrixconfusion_matrix(y_test, preds)Shows which classes get mixed up.
Full reportclassification_report(y_test, preds)Precision, recall and F1 per class.
Cross-validatecross_val_score(model, X, y, cv=5)A far more honest estimate than one split.
Chain the stepsmake_pipeline(StandardScaler(), model)Stops preprocessing leaking across the split.

Every tutorial, by topic

Every scikit-learn tutorial

This is the newest hub on the site and the smallest, so it is also where new tutorials land first. The machine learning hub below covers the theory these models rest on.

Models and training

4

Logistic regression for classification, gradient descent as the optimiser underneath many of these models, genetic algorithms for feature selection, and handling relationships that are not linear.

More scikit-learn tutorials

3

Measuring a model once it is trained — accuracy scores and confusion matrices — plus interview questions.

Keep going

What to learn next to it

Scikit-learn sits at the end of a data pipeline. These cover the rest of it.

Questions people ask

Frequently asked questions

Is scikit-learn enough, or do I need deep learning?

For tabular data — anything that looks like a spreadsheet or a database table — scikit-learn is usually both easier and more accurate, especially gradient boosted trees. Move to Keras or PyTorch when your data is images, audio or free text.

Why is it installed as scikit-learn but imported as sklearn?

Historical naming. The distribution on PyPI is scikit-learn and the Python package inside it is sklearn, so you run pip install scikit-learn and then write from sklearn.... Installing a package literally called sklearn is not what you want.

Do I need to scale my features?

For anything distance-based or regularised — k-nearest neighbours, SVMs, logistic regression, neural networks — yes, or the largest-valued column dominates. Tree models such as random forests and gradient boosting are unaffected by scale, so it is optional there.

What is the difference between fit, transform and fit_transform?

fit learns the parameters, transform applies them, and fit_transform does both. The rule that matters: use fit_transform on the training set and transform on the test set, so nothing about the test data influences the model.

Why is my accuracy 95% but the model is useless?

Almost certainly imbalanced classes. If 95% of rows are one class, predicting that class every time scores 95%. Look at a confusion matrix and the per-class precision and recall instead, and consider stratify=y when splitting.

How do I choose between models?

Start with a simple baseline — logistic regression or a small decision tree — then compare a random forest and gradient boosting using cross_val_score rather than a single split. Spend the time you save on better features, which usually matter more than the choice of algorithm.

Fit one model and score it honestly

Split the data, fit a logistic regression, read the confusion matrix. That is a complete machine learning project in about eight lines.