DSLR
Sorts Hogwarts students into one of four houses from their course scores, with one-vs-all logistic regression and gradient descent written from scratch — no ML library.
TL;DR
- 4 houses
- 13 course features
My partReimplemented a pandas-style describe from scratch: count, mean, standard deviation, quartiles with interpolation, variance, skewness and kurtosis — no NumPy or pandas.
- ROLE
- Solo
- CONTEXT
- École 42 · AI projects
- STATUS
- Done
- STACK
- PythonMatplotlib
- RESULT
- On the 1,600 training students the model reproduces the house labels with 97.9 % accuracy (34 errors). This is a training-set score, measured on the same data it was trained on — there is no held-out test set.
- LAST UPDATED
- 26 Sep 2026
1
Data
Hogwarts student records: 1,600 labelled training students, 13 course scores, 4 houses; the official 400-student test set is unlabelled.
2
What I built
Solo
- Reimplemented a pandas-style describe from scratch: count, mean, standard deviation, quartiles with interpolation, variance, skewness and kurtosis — no NumPy or pandas.
- Implemented logistic regression from scratch: sigmoid, log-loss, gradient and z-score standardisation, with batch, stochastic and mini-batch gradient descent as interchangeable variants.
- Built the one-vs-all multiclass scheme: one binary classifier per house, prediction by argmax over the four sigmoid probabilities, weights exported to JSON.
- Wrote the three exploratory visualisations by hand with Matplotlib (histograms, scatter plots, a coloured pair plot) for feature selection.
3
Key choices
- One-vs-all logistic regression
- Four independent binary classifiers, one per house, and a student is assigned the house with the highest sigmoid probability.
- Three selectable gradient-descent variants
- Batch, stochastic and mini-batch, chosen at runtime, so the optimisation strategies can be compared.
- Manual z-score standardisation
- Every feature is standardised by hand before training and prediction, so scores on very different scales stay comparable.
4
Results
On the 1,600 training students the model reproduces the house labels with 97.9 % accuracy (34 errors). This is a training-set score, measured on the same data it was trained on — there is no held-out test set.
5
Limits
- No held-out evaluation: the official test set is unlabelled, so the only accuracy is measured on the training set itself — not a generalisation score.
- No regularisation and no early stopping; a fixed learning rate and hand-tuned iteration counts.
- The shipped model uses all thirteen features; the feature-selection code is left commented out.
- Minimal cleaning: non-numeric or missing values are replaced by zero rather than imputed.
Questions about this project? → Email me