RL-Collection-v1 is a large-scale, curated corpus for reinforcement learning from verifiable rewards (RLVR) of reasoning-oriented language models. The dataset combines, filters, normalizes, and deduplicates a broad set of public RL datasets into a single consistent schema. It was created by ahmad21omar and last updated on May 20, 2026.
Use Cases
- Training reinforcement learning agents based on machine-verifiable ground-truth signals like math equivalence.
- Benchmarking model reasoning capabilities based on tasks such as code execution or Prolog rule induction.
- Developing reward models for language models based on the corpus's schema validation and multiple-choice signals.
Strengths
- The corpus is described as large-scale and has been curated, filtered, normalized, and deduplicated.
- Each row carries a machine-verifiable ground-truth signal, providing a consistent evaluation mechanism.
- It unifies a broad set of public RL datasets into a single consistent schema.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count, file formats, and license information are unknown, which may limit suitability assessment.
- Data may reflect source bias inherent to the aggregated public RL datasets.
Provenance
- Source
- Aggregated from a broad set of public RL datasets.
- Collection Method
- Combined, filtered, normalized, and deduplicated by the author.
- Time Range
- null
- Freshness
- Last updated 2026-05-20 11:21:54; freshness should be verified.
- Geography
- null