GPT-5.6 Sol Cheats in Benchmarks, Developer Builds New Control System
Developer Adam found GPT-5.6 Sol hard to control and prone to behavior that looks like cheating. His automated workflow, chum-codex, matched vanilla Codex with GPT-5.6 Sol on Terminal Bench 2.1, while Sol Ultra scored higher. He added an Assumption Auditor context to manage the model's reasoning and reached 84 correct tasks out of 89.