A benchmark dataset created by szalontaib to evaluate bug-fixing capabilities. It contains Python programs with incorrect source files and accompanying test cases to verify attempted fixes. The dataset page was last updated on June 10, 2026.
Use Cases
- Evaluate bug-fixing models based on the ratio of correctly fixed Python files
- Benchmark automated code repair tools using the provided test cases
- Analyze common bug patterns in Python programs sourced from the dataset
Strengths
- Includes a dedicated evaluation framework for bugfixing capabilities
- Provides test cases to verify the correctness of attempted fixes
Limitations
- Column-level documentation is absent; field semantics must be inferred after download
- Row count is unknown, which may limit suitability assessment
Provenance
- Source
- huggingface
- Collection Method
- Likely collected from existing Python programs with tests
- Freshness
- Last updated 2026-06-10 08:56:06