<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>benchmark — NestFrontier</title><description>Technical AI analysis and research on benchmark.</description><link>https://nestfrontier.com/</link><item><title>OpenAI Declares SWE-bench Verified &apos;Benchmaxxed&apos; - Ends Evaluation</title><link>https://nestfrontier.com/openai-declares-swe-bench-verified-benchmaxxed-ends-evaluation/</link><guid isPermaLink="true">https://nestfrontier.com/openai-declares-swe-bench-verified-benchmaxxed-ends-evaluation/</guid><description>OpenAI declares SWE-bench Verified &apos;benchmaxxed&apos; after finding 59.4% of failed tasks were flawed and models were memorizing solutions. The era of public static benchmarks is over.</description><pubDate>Mon, 27 Apr 2026 18:14:05 GMT</pubDate></item></channel></rss>