After Dorfman & Ghosh, Developing Games That Learn (1996)

Single-Trial Learner

The O player starts knowing only the rules and picks moves at random. Each time it loses, it walks back from the loss to its point of no return, the last position where it still had a choice, and remembers that position as never-enter. One loss is enough to learn each lesson.

X opponent (moves first)O learner

Progress

0
Games
0
O losses
0
Draws
0
O wins
0
Remembered

Last lesson

Never-enter positions

Each board is a position the learner will refuse to move into. Solid border: the point of no return found right after a loss. Dashed: a position added when the blame walked further back, because every choice from there was already known to lose.

What happens after a loss

  1. Take the position just before X's winning move. X had a win from there, so it's lost. Mark it never-enter.
  2. Step back to the position where O chose the move that led there. List every move O could have made.
  3. If every one of those moves leads to a position already marked, O had no real choice. That position was lost too, so mark the position O's previous move created and step back again.
  4. Stop at the first decision where O had a move not yet known to lose. That move was the mistake. The blame never goes further back than one loss proves.

Positions are stored as board layouts, not move sequences, so the learner avoids a losing position however it would get there. Every mark is a true forced loss, so the learner never throws away a good move. Against perfect play it ends up never losing. This is a reconstruction from the book's published description, not the authors' code. It is the small-game version of the learning in PAX 2.0, where chess positions rarely repeat and blame is harder to place.