Russian-language text pairs for training cross-encoder reranker models. The dataset contains 3.18 million training pairs, composed of 2.39 million examples from over 20 Russian NLI, QA, and paraphrase datasets, 606k TF-IDF hard negatives, and 222k translated MS MARCO entries. It was created by ARGA100 and last updated on June 29, 2026.
Use Cases
- Training cross-encoder rerankers for Russian search engines based on the described NLI/QA/paraphrase data.
- Generating hard negative examples for contrastive learning based on the TF-IDF method mentioned.
- Improving Russian question-answering systems using the translated MS MARCO component.
- Fine-tuning models for Russian paraphrase detection based on the included paraphrase datasets.
Strengths
- Contains 3.18 million total training pairs.
- Aggregates data from over 20 distinct Russian-language source datasets.
- Includes 606k TF-IDF hard negatives to improve model discrimination.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment.
- Last updated 2026-06-29 21:27:58; freshness should be verified.
Provenance
- Source
- ARGA100 on Hugging Face.
- Collection Method
- Aggregated from over 20 Russian NLI/QA/paraphrase datasets, with added TF-IDF hard negatives and a translated MS MARCO subset.
- Freshness
- Last updated 2026-06-29 21:27:58.
- Geography
- Russian-language text.