Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
A fixed evaluation-set composition built with BenchPress, last updated July 7, 2026. The dataset is a mix of established AI benchmarks, sourcing data from repositories like MMLU, ARC Challenge, and HellaSwag at load time. It was created by user seonghyeon0408 and focuses on abilities like evidence retrieval and abstract pattern induction.
The repository stores only a recipe (manifest.json); the actual data is streamed from source repositories at load time, which may affect availability.