AI

Why we no longer evaluate SWE-bench Verified

OpenAI has stopped evaluating the SWE-bench Verified benchmark due to contamination and mismeasurement issues. Analysis found flaws in the testing process, including training leakage, which can lead to inaccurate results. This affects the assessment of progress in coding abilities. The company recommends using SWE-bench Pro instead.
OpenAI has stopped evaluating the SWE-bench Verified benchmark due to contamination and mismeasurement issues. Analysis found flaws in the testing process, including training leakage, which can lead to inaccurate results. This affects the assessment of progress in coding abilities. The company recommends using SWE-bench Pro instead. --- Why it matters: This matters because accurate benchmarks are crucial for evaluating AI's coding capabilities and measuring progress in this area. Inaccurate measurements can hinder research and development efforts. Source: https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified

This article was originally published at: https://openai.com/index/why-we-no-longer-evaluate-swe-be...