Diplomacy RL — Qwen3-235B LoRA

Rank-32 LoRA on Qwen/Qwen3-235B-A22B-Instruct-2507, trained with GRPO to play France in Diplomacy against 6 Grok-4-fast opponents. Fresh LoRA from base — no SFT. This repo is the final adapter, optimizer step 25.

Usage

from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = AutoModelForCausalLM.from_pretrained(
    "Qwen/Qwen3-235B-A22B-Instruct-2507", device_map="auto", torch_dtype="auto"
)
model = PeftModel.from_pretrained(base, "GoodStartLabs/diplomacy-235b-lora")
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-235B-A22B-Instruct-2507")

The policy emits text-format tool calls, not native function calls — <tool_call>{"name": ..., "args": {...}}</tool_call>. It expects the France system prompt and 12-tool schema from training.

Training

  • 25 optimizer updates · 3,184 episodes · 95 h · Dec 2025
  • G16 × B8 = 128 episodes/update · constant LR 2e-4 · importance sampling · temp 1.0 · no KL penalty
  • 1,024-token generation cap · 32,000-token trajectory cap · 10 game phases · 40 turns/phase
  • Reward: 1.0 × Δcenters + 0.2 × (centers − 3) per phase, +0.02/msg in Movement, −0.05/msg in Retreat & Adjustment, −0.1 malformed tool call, −0.5 trajectory overflow
Step reward centers entropy KL turns/ep
0 2.20 3.93 0.308 0.0118 117
6 4.39 4.29 0.346 0.0147 140
12 6.00 4.68 0.352 0.0147 170
18 6.42 4.81 0.481 0.0089 181
24 7.01 4.87 0.355 0.0054 221

Liftoff by updates 3–6, then a staircase. Entropy rose into the middle of the run (peak 0.535 at step 17) and came back. Sampled KL peaked at step 10 and then fell.

Full per-update metrics: training_metrics.csv. The run stopped at 25 of 240 configured updates and had no held-out evaluator, so the numbers above are training reward, not a benchmark.

Environment

GoodStartLabs/rl_envsenvs/tinker_diplomacy/, commit 0c9ef35.

12 tools — submit_all_orders, send_message, check_messages, finish_phase, get_possible_orders, get_game_state, list_units, a persistent diary (write_diary/read_diary), and runtime rulebook access (list_rule_files/cat_rule_file). Opponents are x-ai/grok-4-fast via OpenRouter, frozen.

Trained on Tinker. Earlier checkpoints from the same run are available on request.

Downloads last month
8
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for GoodStartLabs/diplomacy-235b-lora

Adapter
(9)
this model