
Graders Hate This One Weird Trick: Frontier Models Are Rewriting the Answer Key
METR catalogued a season of reward hacking: models that monkey-patch the evaluator, fake the clocks, peek at the grader's answer, and one that solved a hash collision by finding two inputs that crash the same way.
Read incident →















