Learn2Slither
A snake that learns to play through trial and error with tabular Q-learning — a Q-table over the head's four-neighbour vision, trained across thousands of self-play sessions.
TL;DR
- 5,000 training sessions
- 294 learned states
My partBuilt the tabular Q-learning agent: a dictionary Q-table, the temporal-difference update and epsilon-greedy action choice with a decaying epsilon.
- ROLE
- Solo
- CONTEXT
- École 42 · AI projects
- STATUS
- Done
- RESULT
- Trained by self-play on a 10×10 board, the Q-table grew to 294 learned states after 5,000 sessions; the reported peak is a snake of 71 cells around session 4,500 — read from the live counter, not a held-out benchmark.
- LAST UPDATED
- 26 Sep 2026
1
What I built
Solo
- Built the tabular Q-learning agent: a dictionary Q-table, the temporal-difference update and epsilon-greedy action choice with a decaying epsilon.
- Designed the vision-based state encoding — a four-character string of the tiles around the head — and the NumPy game and board engine.
- Set the reward scheme (+10 green apple, -10 red, -1 per move, -100 on death) and the per-session training loop with model save and load.
- Wrote the argparse command-line tool and the pygame GUI (step-by-step mode, speed control, live stats and a learning-curve plot).
2
Key choices
- Tabular Q-learning, not deep RL
- The state is a four-symbol string, so only a few hundred states ever appear (294 after 5,000 sessions) and a plain Q-table suffices — no neural network.
- Head-local vision as the state
- The agent sees only the four tiles next to the head, keeping the state space tiny and learnable; the full line of sight is computed but kept for display.
- Shaped rewards with a heavy death penalty
- -100 on collision against -1 per step pushes the policy toward survival first, then toward the green apples.
3
Results
Trained by self-play on a 10×10 board, the Q-table grew to 294 learned states after 5,000 sessions; the reported peak is a snake of 71 cells around session 4,500 — read from the live counter, not a held-out benchmark.
4
Limits
- The state is only the adjacent cell in each direction, so the agent cannot see apples or walls farther than one tile — it learns collision avoidance more than goal-seeking.
- Performance is high-variance: the reported max length climbs into the 60s-70s then falls back, and per-session reward stays negative throughout.
- Fixed hyperparameters, a single 10×10 board, and single-run maxima rather than an average over many games.
- The max length is read from the live counter, not stored in the saved models, so it is reported rather than reproducible after the fact.
Questions about this project? → Email me