--- dataset_info: features: - name: repo dtype: string - name: instance_id dtype: string - name: base_commit dtype: string - name: patch dtype: string - name: test_patch dtype: string - name: problem_statement dtype: string - name: hints_text dtype: string - name: created_at dtype: string - name: version dtype: string - name: FAIL_TO_PASS dtype: string - name: PASS_TO_PASS dtype: string - name: environment_setup_commit dtype: string splits: - name: dev num_bytes: 4783179 num_examples: 225 - name: test num_bytes: 44121927 num_examples: 2294 - name: train num_bytes: 367610377 num_examples: 19008 download_size: 120086340 dataset_size: 416515483 configs: - config_name: default data_files: - split: dev path: data/dev-* - split: test path: data/test-* - split: train path: data/train-* --- ### Dataset Summary SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 2,294 Issue-Pull Request pairs from 12 popular Python repositories. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The dataset was released as part of [SWE-bench: Can Language Models Resolve Real-World GitHub Issues?](https://arxiv.org/abs/2310.06770) ## Want to run inference now? This dataset only contains the problem_statement (i.e. issue text) and the base_commit which can represents the state of the codebase before the issue has been resolved. If you want to run inference using the "Oracle" or BM25 retrieval settings mentioned in the paper, consider the following datasets. [princeton-nlp/SWE-bench_oracle](https://huggingface.co/datasets/princeton-nlp/SWE-bench_oracle) [princeton-nlp/SWE-bench_bm25_13K](https://huggingface.co/datasets/princeton-nlp/SWE-bench_bm25_13K) [princeton-nlp/SWE-bench_bm25_27K](https://huggingface.co/datasets/princeton-nlp/SWE-bench_bm25_27K) [princeton-nlp/SWE-bench_bm25_40K](https://huggingface.co/datasets/princeton-nlp/SWE-bench_bm25_40K) [princeton-nlp/SWE-bench_bm25_50k_llama](https://huggingface.co/datasets/princeton-nlp/SWE-bench_bm25_50k_llama) ### Supported Tasks and Leaderboards SWE-bench proposes a new task: issue resolution provided a full repository and GitHub issue. The leaderboard can be found at www.swebench.com ### Languages The text of the dataset is primarily English, but we make no effort to filter or otherwise clean based on language type. ## Dataset Structure ### Data Instances An example of a SWE-bench datum is as follows: instance_id: (str) - A formatted instance identifier, usually as repo_owner__repo_name-PR-number. patch: (str) - The gold patch, the patch generated by the PR (minus test-related code), that resolved the issue. repo: (str) - The repository owner/name identifier from GitHub. base_commit: (str) - The commit hash of the repository representing the HEAD of the repository before the solution PR is applied. hints_text: (str) - Comments made on the issue prior to the creation of the solution PR’s first commit creation date. created_at: (str) - The creation date of the pull request. test_patch: (str) - A test-file patch that was contributed by the solution PR. problem_statement: (str) - The issue title and body. version: (str) - Installation version to use for running evaluation. environment_setup_commit: (str) - commit hash to use for environment setup and installation. FAIL_TO_PASS: (str) - A json list of strings that represent the set of tests resolved by the PR and tied to the issue resolution. PASS_TO_PASS: (str) - A json list of strings that represent tests that should pass before and after the PR application. [More Information needed](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)Follow this markdown format, put it so i can copy pasteEditmarkdown--- dataset_info: features: - name: repo dtype: string - name: instance_id dtype: string - name: base_commit dtype: string - name: patch dtype: string - name: test_patch dtype: string - name: problem_statement dtype: string - name: hints_text dtype: string - name: created_at dtype: string - name: version dtype: string - name: FAIL_TO_PASS dtype: string - name: PASS_TO_PASS dtype: string - name: environment_setup_commit dtype: string splits: - name: dev num_bytes: 4783179 num_examples: 225 - name: test num_bytes: 44121927 num_examples: 2294 - name: train num_bytes: 367610377 num_examples: 19008 download_size: 120086340 dataset_size: 416515483 configs: - config_name: default data_files: - split: test path: data/test-* --- # Dataset Card for SWE-bench ## Dataset Summary SWE-bench Pro is a dataset that tests systems' ability to solve GitHub issues automatically. This dataset is the public set only. The dataset was released as part of SWE-Bench Pro paper. ## Dataset Details ### Dataset Description - **Curated by:** Scale AI - **Language(s) (NLP):** English (primarily) ## Uses ### Direct Use This dataset is designed for evaluating systems' ability to automatically resolve GitHub issues by generating appropriate code patches. ### Out-of-Scope Use This dataset is not meant as training data. ### Data Instances An example of a SWE-bench datum is as follows: ```json { "instance_id": "repo_owner__repo_name-PR-number", "patch": "The gold patch that resolved the issue", "repo": "The repository owner/name identifier from GitHub", "base_commit": "The commit hash before the solution PR is applied", "hints_text": "Comments made on the issue prior to the solution PR", "created_at": "The creation date of the pull request", "test_patch": "A test-file patch contributed by the solution PR", "problem_statement": "The issue title and body", "version": "Installation version for running evaluation", "environment_setup_commit": "Commit hash for environment setup and installation", "FAIL_TO_PASS": "JSON list of tests resolved by the PR", "PASS_TO_PASS": "JSON list of tests that should pass before and after PR" }