Afficher le post originalMasquer le post original
A large share of the code on the 3D Godot game I'm working on is written by Claude Code. The agent can't watch the game, so we gave it a way to check its own work from images and logs. It writes a throwaway driver script that plays a scripted sequence, renders every frame with Godot's Movie Maker mode, reads the log, and then reviews the footage as a contact sheet followed by full-resolution frames.
That worked until the game hid a failure from it. A clip re-export made every animation one frame shorter, the animation loader's guard switched off all clip animation, and the game fell back to a procedural walk, which is what it's designed to do. The character still walked, just stiffly. Claude passed that build three separate times across three sessions, and the warning that explained it was in every render log that day. I noticed the next day because I know how the character is supposed to move.
What changed afterwards:
• The log gets searched for WARNING before any frame is opened. If a warning names the system being tested, the footage shows the fallback and there's nothing to review.
• Drivers have to confirm the system under test is actually running, not just that their measurements look right. A driver that only checks its own numbers can pass while filming the fallback.
• Evidence from a branch expires when the branch merges. One of the three passes was a render made on a branch cut before the break.
The biggest addition is a second reviewer agent that only ever sees pictures. The session running the test has to give it four things or it refuses: the clip, the choreography with timestamps, what correct looks like, and which fallback could be hiding the problem. It doesn't see logs or code. It lists findings with the frames it's basing them on and has to commit to one of two verdicts, "looks correct" or "something is wrong, and here's what." It isn't allowed to answer an anomaly it can't explain with a request for a better render.
That last rule came from a comparison we ran three days later. The agent re-rendered the broken build and a healthy one, and blind reviewers on two Claude models each got the same cropped frames with no logs. On the broken build, Fable 5 said "something is wrong with this build, specifically the weapon" with high confidence. Opus 5 saw the same anomalies, suggested occlusion as the explanation, and concluded nothing indicated a broken build. That's one sample from the August 2026 models, so read it as an anecdote. It's still why the reviewer agent is pinned to Fable and required to commit.
The write-up has the whole loop and a couple of other ways a test driver can give you a confident wrong answer, including one that failed on every run for a bug that never existed: https://protoforgesystems.com/devlog/posts/how-a-coding-agent-verifies-a-3d-game
The driver harness, the frame-review scripts and the reviewer agent are an MIT-licensed Claude Code plugin: https://github.com/ProtoForgeSystems/protoforge-claude-plugin-game-review
/plugin marketplace add ProtoForgeSystems/protoforge-claude-plugin-game-review /plugin install game-review@protoforge-game-review