Add eval.yaml (#4)
- Add eval.yaml (0bffd3f4825f713e220b52e99eb3209fd98b4162) Co-authored-by: Niels Rogge <nielsr@users.noreply.huggingface.co>
This commit is contained in:
committed by
system
co-authored by
Niels Rogge
parent
2dd05cab15
commit
7ab5114912
@@ -0,0 +1,8 @@
|
|||||||
|
name: SWE-Bench Pro
|
||||||
|
description: SWE-Bench Pro is an advanced, 1,865-task benchmark for evaluating AI coding agents on realistic, complex, long-horizon software engineering tasks. It improves upon the original SWE-Bench by focusing on non-trivial patches across 41+ repositories, aiming to address data contamination, task and complexity.
|
||||||
|
evaluation_framework: swe-bench-pro
|
||||||
|
|
||||||
|
tasks:
|
||||||
|
- id: SWE_Bench_Pro
|
||||||
|
config: default
|
||||||
|
split: test
|
||||||
Reference in New Issue
Block a user