01
Set up a fair comparison
- Give Claude, ChatGPT and Gemini the same project, bug and exact prompt. Start every attempt from the same clean copy.
- Back up your tests somewhere the AI cannot edit them. Restore that copy after each attempt and run the tests yourself.
- Run each AI at least twice. Keep the model, tool settings and starting files consistent across the comparison.
02
The task to copy
Fix this bug: [what goes wrong, and the exact error message].
The project is in [folder].
Don't change, delete or skip any test.
When you're done, list every file you changed.
04
Pick the winner
- A run qualifies only when the bug is fixed, the original tests pass, and no test was edited, deleted or skipped.
- Compare time and cost among the AIs that qualify in every run. Record actual run charges when available; do not invent a per-run cost for a flat subscription.