Opper Releases Jevman, A Pac-Man Benchmark For Decision Models
Opper published jevman, an open-source benchmark that has AI models play Pac-Man to compare decision-making. Six models, including Jev, Kev, Laya, Clef, and GPT-6 Luna Decisions, each play 100 games against standard-rule ghosts. Models receive maze, dot, and ghost data as JSON and return up, down, left, or right as probabilities within a 2-second limit per junction. Ranking uses the 100-game average score, with a 95 percent confidence margin, so models inside that margin are treated as tied.