Sign in to view source links and access this dataset
Description
LitRetrieval is a dataset for training and evaluating text retrieval, sentence similarity, and instruction-following models on literary texts in Russian and English. The dataset features deep semantic challenges, complex narrative vocabulary, and metaphorical structures typical of classic and contemporary literature. It was created by ImpulseLeap and last updated on July 16, 2026.
Use Cases
Training text retrieval models based on rich literary texts in Russian and English.
Evaluating sentence similarity models based on complex narrative vocabulary and metaphorical structures.
Benchmarking instruction-following models on semantic challenges present in classic and contemporary literature.
Strengths
Covers two languages, Russian and English, enabling multilingual model development.
Focuses on literary texts, which likely contain complex narrative vocabulary and metaphorical structures.
Designed for specific NLP tasks: text retrieval, sentence similarity, and instruction-following.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count is unknown, which may limit suitability assessment.
Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
Source
ImpulseLeap on Hugging Face.
Freshness
Last updated 2026-07-16 18:47:53; freshness should be verified.
License is unknown; terms of use must be verified before application.