The O player starts knowing only the rules and picks moves at random. Each time it loses, it walks back from the loss to its point of no return, the last position where it still had a choice, and remembers that position as never-enter. One loss is enough to learn each lesson.
Each board is a position the learner will refuse to move into. Solid border: the point of no return found right after a loss. Dashed: a position added when the blame walked further back, because every choice from there was already known to lose.
Positions are stored as board layouts, not move sequences, so the learner avoids a losing position however it would get there. Every mark is a true forced loss, so the learner never throws away a good move. Against perfect play it ends up never losing. This is a reconstruction from the book's published description, not the authors' code. It is the small-game version of the learning in PAX 2.0, where chess positions rarely repeat and blame is harder to place.