An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's new MirrorCode benchmark tests whether AI models can recreate entire programs on their own. Claude Opus 4. 7 leads with 56 percent, but every model still fails on the most complex tasks.
Read on The-decoder ↗