Nudging a one-move search, from easy and hard starting boards
rejectedevidence: task-record onlypublic informationA learned correction bolted onto a simple search, trained half on fresh games and half on difficult mid-game positions — it improved on the search slightly and stayed far below the reference.
A learned correction bolted onto a simple search, trained half on fresh games and half on difficult mid-game positions.
The intuition
Two ideas that recur across this family, combined.
Learn a correction, not a policy. Starting a policy network from nothing means spending most of the budget rediscovering that filling a column to the ceiling is bad. Instead, take an exact one-move fair search, normalise its opinion of each column into scores, and let a small network add a correction on top. Set that correction to zero and the policy is bit-for-bit identical to the search. Every point of difference is therefore attributable to what was learned, and there is a safe fallback by construction.
Do not only start from the beginning. A policy trained solely from opening boards sees crowded, dangerous mid-game positions only after it has learned to survive long enough to reach them — which is exactly the skill it lacks. So half of every training round starts from a restart state: a public position extracted from a long trajectory, replayed forward with a fresh, independent random future.
The network is also built to be exactly mirror-equivariant, so it cannot learn an arbitrary left-right preference from the noise in its data.
How it works, step by step
- Score the columns exactly. Run the fixed fair one-move search and normalise its per-column values.
- Add a learned residual. A small two-pass network, structurally guaranteed to be reflection-equivariant, adjusts those scores.
- Collect experience. Each iteration plays 512 episodes: 256 from a fresh opening board with a long move cap, 256 from a curriculum restart position with a 100-move horizon.
- Update by policy gradient. Four passes over each batch, minibatches of 512, across 64 iterations — 32,768 episodes in total.
- Gate. A 32-game screen stands between training and any further stage.
What happened
It modestly improved on the one-move search it was correcting, and remained far below both the sustainable-flow targets it was given and the depth-4 reference. That sentence is the whole retained result.
The technical record
Status in the experiment index: rejected,
task-record only. docs/exploratory/audit-04-blind-spots.md lists
curriculum-option-ppo.cpp among 30 sources whose rejection "exists only as a
task record, with no ledger protocol, no artifact hash, and no retained per-game
data". No result number exists to quote.
The configuration facts above come from the source and from
docs/exploratory/audit-05-optimistic-curriculum.md, which records the mixing
rule as a fixed 50/50 split of initial and mature-restart episodes per
iteration, in a single arm — the ratio was never ablated, so nothing is
known about whether the curriculum helped, hurt, or did nothing.
From the source: training lane 0x3d670000...0x3d678000, Stage-A screen
0x3d680000...0x3d680020 (32 games), 64 iterations of 512 episodes, four update
epochs, minibatch 512, a 1,000-move cap on initial episodes and a 100-move
curriculum horizon, and 4,096 expected curriculum states. The source reuses the
viability-controller program's public-state representation.
Source: curriculum-option-ppo.cpp.
What this taught us, and what is still open
A safe-fallback residual is a good pattern, and it recurs. Several later programs in the repository adopted the same structure — an exact search plus a zero-initialised learned correction — because it makes "did learning help?" a well-posed question with a trivially available control. This experiment is an early instance.
The curriculum question is genuinely unanswered. Mixing restart states into
training appears in at least four programs across this repository, always at a
fixed ratio and always as a single arm. docs/exploratory/audit-05-optimistic-curriculum.md
concludes that the order and proportion of oracle-derived and ordinary data
were never ablated as independent factors — not once. So no page here, including
this one, can say whether the curriculum was doing anything.
Still open, and cheap. Because the residual reduces exactly to the base search at zero, a properly powered comparison of "residual, curriculum on" and "residual, curriculum off" against the unmodified search is a small, well-defined experiment that has never been run.
Source files
README.mdxcurriculum-option-ppo.cpp