SLM-Bench is a programmatic benchmark containing 3,000 questions designed to evaluate language models with fewer than 10 million parameters. It covers six core capability areas with 500 questions each, including arithmetic problems. The dataset was created by liodon-ai and was last updated on June 14, 2026.
Use Cases
- Benchmarking arithmetic reasoning capabilities based on the 500 math problems with plausible distractors.
- Evaluating model weaknesses across six core capability areas based on the granular insights the benchmark provides.
- Comparing performance of sub-10M parameter models based on the standardized, programmatic test suite.
Strengths
- Contains 3,000 total questions for evaluation.
- Includes 500 questions specifically for arithmetic capability.
- Designed specifically for sub-10M parameter language models.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment.
- Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
- Source
- liodon-ai
- Collection Method
- Programmatically generated benchmark.
- Freshness
- Last updated 2026-06-14 01:59:57; freshness should be verified.