Skip to content

DSLR

Sorts Hogwarts students into one of four houses from their course scores, with one-vs-all logistic regression and gradient descent written from scratch — no ML library.

TL;DR
  • 4 houses
  • 13 course features

My partReimplemented a pandas-style describe from scratch: count, mean, standard deviation, quartiles with interpolation, variance, skewness and kurtosis — no NumPy or pandas.

ROLE
Solo
CONTEXT
École 42 · AI projects
STATUS
Done
RESULT
On the 1,600 training students the model reproduces the house labels with 97.9 % accuracy (34 errors). This is a training-set score, measured on the same data it was trained on — there is no held-out test set.
LAST UPDATED
26 Sep 2026
scores13 featuresz-scoreby handlogistic reg.one-vs-all ×4argmax4 probabilitieshouseprediction
Fig. — how it works
1

Data

Hogwarts student records: 1,600 labelled training students, 13 course scores, 4 houses; the official 400-student test set is unlabelled.

2

What I built

Solo

  1. Reimplemented a pandas-style describe from scratch: count, mean, standard deviation, quartiles with interpolation, variance, skewness and kurtosis — no NumPy or pandas.
  2. Implemented logistic regression from scratch: sigmoid, log-loss, gradient and z-score standardisation, with batch, stochastic and mini-batch gradient descent as interchangeable variants.
  3. Built the one-vs-all multiclass scheme: one binary classifier per house, prediction by argmax over the four sigmoid probabilities, weights exported to JSON.
  4. Wrote the three exploratory visualisations by hand with Matplotlib (histograms, scatter plots, a coloured pair plot) for feature selection.
3

Key choices

One-vs-all logistic regression
Four independent binary classifiers, one per house, and a student is assigned the house with the highest sigmoid probability.
Three selectable gradient-descent variants
Batch, stochastic and mini-batch, chosen at runtime, so the optimisation strategies can be compared.
Manual z-score standardisation
Every feature is standardised by hand before training and prediction, so scores on very different scales stay comparable.
4

Results

On the 1,600 training students the model reproduces the house labels with 97.9 % accuracy (34 errors). This is a training-set score, measured on the same data it was trained on — there is no held-out test set.

5

Limits

  • No held-out evaluation: the official test set is unlabelled, so the only accuracy is measured on the training set itself — not a generalisation score.
  • No regularisation and no early stopping; a fixed learning rate and hand-tuned iteration counts.
  • The shipped model uses all thirteen features; the feature-selection code is left commented out.
  • Minimal cleaning: non-numeric or missing values are replaced by zero rather than imputed.

Questions about this project? → Email me