Jeff
Loading… GPU — —Training
No training runs yet.
Eval loss
Dev-set loss (NLL) before calibration — lower is better
Training loss
Per update, 10-update moving average
Dev accuracy
Higher is better
All runs
Panel results
Accuracy by model
Each of our models at its best (trained on public + synthetic data) next to the published models. Colour = model family, hatched = the smaller model of a family, grey = published. 🏆 marks the best model on each chart, published ones included. Same 0–100% axis everywhere; untrained and public-only results are in the table below.
Speed
Time per decision
200 panel questions (40 per benchmark), timed from raw text to probabilities: tokenising, one forward pass and the readout.
"One at a time" is what a single API caller waits for; "batched" is throughput when many decisions are sent together.
Games
Zero-shot game tests
No game in training. Each turn the model gets the game state as text and picks one option; the
wording says how the options are phrased. Score: Doom kills per episode (defend_the_center, seed 1234), Frogger crossings
(3 lives, 300 turns), Pac-Man pellets eaten (of 98; 3 lives, 600 turns). Mean ± standard deviation over the episodes.
Synthetic generation
No generation runs yet.