Agents score 21% on Python tasks and 4% when C or C++ is involved. That gap is the whole reason for this project.
I am building a Game Boy emulator in C++, driven by AI coding agents, and publishing the result against the public test ROMs. The pass count is the scoreboard, and it starts at zero.
Every claim here ships with its artefact: the spec, the diff, the commit, the test output. Including the failures.