Small models.
Measured decisions.
Two GPU training experiments. One shared dataset. Testable results.
v0.2 · Play live · Inspect real model decisions
Watch it play. Take the controls.
Loading our trained tiny model…
The downloaded network scores current game-state features on your CPU. Switch to “I play” to take over; Reset starts a fresh game. This is a new browser run, separate from the DGX benchmark below.
What do 4,096, 512, 50.2% and 596M mean?
4,096 synthetic decision examples were used for training. 512 separate states were used for validation and checkpoint selection. The tiny model chose the search teacher’s preferred move in 257 of 512 states (50.2%). This is decision agreement, not game win rate or the percentage of pellets collected.
596M = 596 million parameters (5.96 亿) in the separate Qwen3-0.6B experiment. LoRA trained about 10.1 million adapter parameters. The live game above uses our independently trained 17,601-parameter network. A parameter is a learned numeric weight; it is not a training example.
This game project uses structured state observations. It does not use Jev-Omni weights and does not process images or video. Our earlier multimodal demo is a separate project.
Full game recording · 65 seconds
Download MP4 ↓Tiny policy · seed 100 · 211 pellets · 2,670 points. One complete browser run, shown from start to game over.
Inspect a decision
Same input: original Qwen, LoRA pilot, compact policy, and teacher.
Inspect the exact state sent to the model
Three held-out games
Seeds 100–102 · realtime clock · 180-second simulated limit · one fixed maze
All individual outcomes are shown. Qwen runs on GPU; the 17.6K policy runs on CPU. These three games are an exploratory pilot.