CHSM8 player model: Wesley So
LegumMagister/chsm8-player-base fine-tuned to play like Wesley So: supervised next-move prediction on Wesley So's own moves only (the opponent's moves are context, not targets).
- Data: 5,300 unique games (chess.com + lichess); held-out val 663 / test 661 games = the player's most recent games (from 2024-05-15), never trained on.
- Recipe (chosen on Nakamura from SFT / LoRA / DPO / SFT->DPO ablations; best held-out move-match within an Elo guard): full fine-tune, AdamW lr 3e-5, cosine, 2 passes over the unique games, batch 32 x 512, 1 GPU.
Results (player's own moves, held-out test games)
| base (chsm8-player-base) | this model | |
|---|---|---|
| move-match (top-1 = move played) | 0.5467 | 0.5642 |
| move-match top-3 | - | 0.7929 |
| cross-entropy on the player's moves | 2.9624 | 2.9686 |
Strength: 1910 (95% CI 1891-1928, 1000 games) vs Stockfish 17.1 UCI_Elo 2000 (0.05 s/move).
Limitations: imitates the moves in the available games (mostly online blitz/rapid for active players, classical OTB for historical ones); it is not the player and does not reproduce their strength.
Usage
The checkpoint is a PyTorch dict (model_state_dict, cfg) for the CHSM8 decoder (factored move head over
(kind, src, dst, piece, promo)). Moves are fed as factored rows, BOS first; there is no text input, and play picks the
best legal move by summed factored log-probability. The model classes and the loading and play code will be released later with the project code.
Model tree for LegumMagister/chsm8-player-so
Base model
LegumMagister/chsm8-pt