Studybench is a benchmark of expert-level coding questions about real open-source codebases. Each question is paired with a gold answer and a weighted, source-grounded grading rubric. The dataset was created by jacobli and was last updated on June 12, 2026.
Use Cases
- Benchmarking code generation models based on expert-level questions about real libraries.
- Training models to produce working code based on specific library/framework requirements.
- Developing automated grading systems for code using source-grounded rubrics.
Strengths
- Questions are grounded in real open-source codebases.
- Each question includes a gold answer and a detailed grading rubric.
- Rubrics decompose answers into discrete, checkable claims tied to exact source lines.
Limitations
- Description metadata is limited; actual data quality requires manual inspection after download.
- Row count is unknown, which may limit suitability assessment.
- Column-level documentation is absent; field semantics must be inferred after download.
Provenance
- Source
- huggingface
- Freshness
- Last updated 2026-06-12 02:18:55; freshness should be verified.