Keep the environment fixed
Use the same task inputs, tool surface, permission scope and success criteria for every model.
Score outcomes, not eloquence
Prefer measurable task results, safe incompleteness and verified real-world outcomes over how convincing the final prose sounds.
Include failure cases
Timeouts, malformed output and unavailable capabilities reveal compatibility differences that happy-path benchmarks hide.
Report limitations
A strong result on one frozen suite does not prove one model is universally better. Publish the environment and scope so the result can be interpreted honestly.