Sign in to view source links and access this dataset
Description
180 self-contained operations research tasks benchmarked from academic literature. Each task includes a natural-language problem description, mathematical formulation, reference Gurobi implementation, test instances, and an automated feasibility checker. Created by SmartOR, the dataset was last updated on June 14, 2026.
Use Cases
Benchmarking LLM performance on translating natural-language OR problems into executable code based on the provided problem descriptions.
Evaluating the correctness of generated optimization code using the automated feasibility checker.
Training or fine-tuning models for mathematical reasoning and code generation tasks based on the structured problem-solution pairs.
Reproducing and comparing optimization algorithms using the provided reference Gurobi implementations and test instances.
Strengths
Contains 180 distinct, literature-grounded tasks, providing a substantial benchmark scale.
Each task is a self-contained reproducible unit with multiple components, including a reference implementation and automated checker.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count is unknown, which may limit suitability assessment.
Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
Source
SmartOR
Collection Method
Packaged from academic literature, likely as described in the benchmark paper.
Freshness
Last updated 2026-06-14 11:13:13.
License is unknown; users must verify permissions before use.