2 comments

  • maxnyriev 34 minutes ago ago

    how many of the "fixes" were to code that wasn't broken? The prompt tells the model there are bugs, and these models are eager to please, so I'd expect plenty of confident edits to working code. The original test suite only catches the ones that break something. Precision next to recall would make the leaderboard a lot more useful.

  • ofirpress 4 hours ago ago

    Hi I'm another one of the co-authors. Our team previously built SWE-bench. We think SWE-Sweep is one of the most exciting research directions to work on right now. Happy to answer questions.