- Add eval.yaml (0bffd3f4825f713e220b52e99eb3209fd98b4162) Co-authored-by: Niels Rogge <nielsr@users.noreply.huggingface.co>
8 lines
440 B
YAML
8 lines
440 B
YAML
name: SWE-Bench Pro
|
|
description: SWE-Bench Pro is an advanced, 1,865-task benchmark for evaluating AI coding agents on realistic, complex, long-horizon software engineering tasks. It improves upon the original SWE-Bench by focusing on non-trivial patches across 41+ repositories, aiming to address data contamination, task and complexity.
|
|
evaluation_framework: swe-bench-pro
|
|
|
|
tasks:
|
|
- id: SWE_Bench_Pro
|
|
config: default
|
|
split: test |