Skip to content

Learn2Slither

A snake that learns to play through trial and error with tabular Q-learning — a Q-table over the head's four-neighbour vision, trained across thousands of self-play sessions.

TL;DR
  • 5,000 training sessions
  • 294 learned states

My partBuilt the tabular Q-learning agent: a dictionary Q-table, the temporal-difference update and epsilon-greedy action choice with a decaying epsilon.

ROLE
Solo
CONTEXT
École 42 · AI projects
STATUS
Done
RESULT
Trained by self-play on a 10×10 board, the Q-table grew to 294 learned states after 5,000 sessions; the reported peak is a snake of 71 cells around session 4,500 — read from the live counter, not a held-out benchmark.
LAST UPDATED
26 Sep 2026
stepTD updatestate4 tiles around headQ-tableε-greedyaction4 movesreward+10 / -10 / -1 / -100new stateboard · numpy
Fig. — how it works
1

What I built

Solo

  1. Built the tabular Q-learning agent: a dictionary Q-table, the temporal-difference update and epsilon-greedy action choice with a decaying epsilon.
  2. Designed the vision-based state encoding — a four-character string of the tiles around the head — and the NumPy game and board engine.
  3. Set the reward scheme (+10 green apple, -10 red, -1 per move, -100 on death) and the per-session training loop with model save and load.
  4. Wrote the argparse command-line tool and the pygame GUI (step-by-step mode, speed control, live stats and a learning-curve plot).
2

Key choices

Tabular Q-learning, not deep RL
The state is a four-symbol string, so only a few hundred states ever appear (294 after 5,000 sessions) and a plain Q-table suffices — no neural network.
Head-local vision as the state
The agent sees only the four tiles next to the head, keeping the state space tiny and learnable; the full line of sight is computed but kept for display.
Shaped rewards with a heavy death penalty
-100 on collision against -1 per step pushes the policy toward survival first, then toward the green apples.
3

Results

Trained by self-play on a 10×10 board, the Q-table grew to 294 learned states after 5,000 sessions; the reported peak is a snake of 71 cells around session 4,500 — read from the live counter, not a held-out benchmark.

4

Limits

  • The state is only the adjacent cell in each direction, so the agent cannot see apples or walls farther than one tile — it learns collision avoidance more than goal-seeking.
  • Performance is high-variance: the reported max length climbs into the 60s-70s then falls back, and per-session reward stays negative throughout.
  • Fixed hyperparameters, a single 10×10 board, and single-run maxima rather than an average over many games.
  • The max length is read from the live counter, not stored in the saved models, so it is reported rather than reproducible after the fact.

Questions about this project? → Email me