Letting Jev play my own game — how the hand-off changes the score
A controlled experiment in which TypeSafe's Jev decision model played Charge Run, my one-key dodging game. Over five rounds I changed how the game state was handed over, and compared Jev with random play and my own hand-written policy on the same 20 courses. What raised the score was the hand-off and code-side safety; under these conditions, Jev's own choices did not show a measurable benefit.
- Type
- Self-initiated
- Role
- Experiment design, build, analysis
- Assets used
- My own game, Jev, OpenJev
- Delivered as
- Four-way comparison video (38 s, no audio)
§ 01 — Problem
Can a game that demands sub-second decisions be handed to an LLM-based decision model? And if it can, is the gain coming from the model's judgment or from the information and scaffolding I give it? Without separating the two, "it worked when I let the AI play" is not evidence of anything.
§ 02 — Decision
I fixed the interface: on every landing, hand over the current state and ask for one of three moves — wait, small jump, big jump. As baselines I used players that ignore the state (always big-jump, random) and my own hand-written policy, and ran everyone on the same 20 courses, comparing medians and per-course wins. Each round changed only one thing about the hand-off, and for all five rounds the pass criterion was written down before running (the final lookahead round was a score-maximizing attempt run outside that protocol). The steps were: raw position and velocity → a code-computed "does this move collide?" → code removes colliding moves first → keep only moves that survive the next 0.5 seconds.


§ 03 — Learning
The hand-off moved the score. Given collisions as a count, the model preferred the moves that collided; given a true/false "collides", it avoided them, and the median went up five-fold (56 → 285). Removing colliding moves in code first was the first setup that beat the always-big-jump baseline, but random choice among the remaining moves did just as well as Jev (11 wins, 9 losses). Adding lookahead raised the score from 726 to 1863 — yet the rule-only policy without Jev reached 1698. Writing "prefer the big jump among safe moves" into the prompt made Jev's scores identical to the always-big-jump player on 14 of 20 courses. Total Jev usage cost was $0.53. Next time I would size the number of courses needed to detect the model's own contribution before running, and handle real-time latency (two frames of delay cut the score to about a quarter) by queueing moves ahead.