OpenAI Declares SWE-bench Verified 'Benchmaxxed' - Ends Evaluation
OpenAI declares SWE-bench Verified 'benchmaxxed' after finding 59.4% of failed tasks were flawed and models were memorizing solutions. The era of public static benchmarks is over.
TOPIC_INDEX
1 published entry in this topic.
OpenAI declares SWE-bench Verified 'benchmaxxed' after finding 59.4% of failed tasks were flawed and models were memorizing solutions. The era of public static benchmarks is over.