OpenAI finds roughly 30 percent of popular AI coding test is broken
OpenAI finds roughly 30 percent of popular AI coding test is broken Key Points - OpenAI is pulling its endorsement of the AI coding test SWE-Bench Pro after a review found roughly 30 percent of its tasks are flawed.
Read on The-decoder ↗