基本信息
- 来源: blogs_podcasts
- 原始来源: https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified
来源摘要/节选
公开展示已截断至最多 800 个字符;请访问原始来源查看完整上下文。
SWE-bench Verified is increasingly contaminated. We recommend SWE-bench Pro.
Since we first published SWE-bench Verified in August 2024, the industry has widely used it to measure the progress of models on autonomous software engineering tasks. After its release, SWE-bench Verified provided a strong signal of capability progress and became a standard metric reported in frontier model releases. Tracking and forecasting progress of these capabilities is also an important part of OpenAI’s Preparedness Framework . When we created the Verified benchmark initially, we attempted to solve issues in the original evaluation that made certain tasks impossible to accomplish in the SWE-bench dataset (opens in a new window) .…
来源说明
当前只保存了公开页面节选,不代表原文全文。请以原始来源为准。
本页只呈现已做哈希绑定的来源证据,不包含基于旧正文或缺失原文的扩展推断。