Files
SWE-bench_Pro/README.md
T
2025-09-19 18:11:33 +00:00

6.1 KiB
Raw Blame History

dataset_info, configs
dataset_info configs
features splits download_size dataset_size
name dtype
repo string
name dtype
instance_id string
name dtype
base_commit string
name dtype
patch string
name dtype
test_patch string
name dtype
problem_statement string
name dtype
hints_text string
name dtype
created_at string
name dtype
version string
name dtype
FAIL_TO_PASS string
name dtype
PASS_TO_PASS string
name dtype
environment_setup_commit string
name num_bytes num_examples
dev 4783179 225
name num_bytes num_examples
test 44121927 2294
name num_bytes num_examples
train 367610377 19008
120086340 416515483
config_name data_files
default
split path
dev data/dev-*
split path
test data/test-*
split path
train data/train-*

Dataset Summary

SWE-bench is a dataset that tests systems ability to solve GitHub issues automatically. The dataset collects 2,294 Issue-Pull Request pairs from 12 popular Python repositories. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Want to run inference now?

This dataset only contains the problem_statement (i.e. issue text) and the base_commit which can represents the state of the codebase before the issue has been resolved. If you want to run inference using the "Oracle" or BM25 retrieval settings mentioned in the paper, consider the following datasets. princeton-nlp/SWE-bench_oracle princeton-nlp/SWE-bench_bm25_13K princeton-nlp/SWE-bench_bm25_27K princeton-nlp/SWE-bench_bm25_40K princeton-nlp/SWE-bench_bm25_50k_llama

Supported Tasks and Leaderboards

SWE-bench proposes a new task: issue resolution provided a full repository and GitHub issue. The leaderboard can be found at www.swebench.com

Languages

The text of the dataset is primarily English, but we make no effort to filter or otherwise clean based on language type.

Dataset Structure

Data Instances

An example of a SWE-bench datum is as follows:

instance_id: (str) - A formatted instance identifier, usually as repo_owner__repo_name-PR-number. patch: (str) - The gold patch, the patch generated by the PR (minus test-related code), that resolved the issue. repo: (str) - The repository owner/name identifier from GitHub. base_commit: (str) - The commit hash of the repository representing the HEAD of the repository before the solution PR is applied. hints_text: (str) - Comments made on the issue prior to the creation of the solution PRs first commit creation date. created_at: (str) - The creation date of the pull request. test_patch: (str) - A test-file patch that was contributed by the solution PR. problem_statement: (str) - The issue title and body. version: (str) - Installation version to use for running evaluation. environment_setup_commit: (str) - commit hash to use for environment setup and installation. FAIL_TO_PASS: (str) - A json list of strings that represent the set of tests resolved by the PR and tied to the issue resolution. PASS_TO_PASS: (str) - A json list of strings that represent tests that should pass before and after the PR application.

More Information neededFollow this markdown format, put it so i can copy pasteEditmarkdown--- dataset_info: features:

  • name: repo dtype: string
  • name: instance_id dtype: string
  • name: base_commit dtype: string
  • name: patch dtype: string
  • name: test_patch dtype: string
  • name: problem_statement dtype: string
  • name: hints_text dtype: string
  • name: created_at dtype: string
  • name: version dtype: string
  • name: FAIL_TO_PASS dtype: string
  • name: PASS_TO_PASS dtype: string
  • name: environment_setup_commit dtype: string splits:
  • name: dev num_bytes: 4783179 num_examples: 225
  • name: test num_bytes: 44121927 num_examples: 2294
  • name: train num_bytes: 367610377 num_examples: 19008 download_size: 120086340 dataset_size: 416515483 configs:
  • config_name: default data_files:
    • split: test path: data/test-*

Dataset Card for SWE-bench

Dataset Summary

SWE-bench Pro is a dataset that tests systems' ability to solve GitHub issues automatically. This dataset is the public set only.

The dataset was released as part of SWE-Bench Pro paper.

Dataset Details

Dataset Description

  • Curated by: Scale AI
  • Language(s) (NLP): English (primarily)

Uses

Direct Use

This dataset is designed for evaluating systems' ability to automatically resolve GitHub issues by generating appropriate code patches.

Out-of-Scope Use

This dataset is not meant as training data.

Data Instances

An example of a SWE-bench datum is as follows:

{
  "instance_id": "repo_owner__repo_name-PR-number",
  "patch": "The gold patch that resolved the issue",
  "repo": "The repository owner/name identifier from GitHub",
  "base_commit": "The commit hash before the solution PR is applied",
  "hints_text": "Comments made on the issue prior to the solution PR",
  "created_at": "The creation date of the pull request",
  "test_patch": "A test-file patch contributed by the solution PR",
  "problem_statement": "The issue title and body",
  "version": "Installation version for running evaluation",
  "environment_setup_commit": "Commit hash for environment setup and installation",
  "FAIL_TO_PASS": "JSON list of tests resolved by the PR",
  "PASS_TO_PASS": "JSON list of tests that should pass before and after PR"
}