Sign in to view source links and access this dataset
Description
220,222 synthetic data points for agricultural advisory tasks, generated by Google Gemini 2.5 Flash. The dataset is designed for instruction tuning and chain-of-thought reasoning, with Hindi as the target output language and English used for internal reasoning and metadata. Soketlabs published the dataset on Hugging Face, with a last update recorded on January 16, 2026.
Use Cases
Train instruction-following models for agricultural advice based on the described task category.
Develop chain-of-thought reasoning models for agricultural problem-solving based on the dataset's design.
Fine-tune language models for Hindi-language agricultural Q&A based on the target output language.
Benchmark synthetic data generation methods for domain-specific tasks based on the use of a source model.
Strengths
Contains 220,222 total data points, providing a substantial volume for training.
Dataset is structured for chain-of-thought and instruction tuning, a specific and useful task format.
Explicitly targets Hindi as the output language, addressing a specific linguistic need.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Description metadata is limited; actual data quality requires manual inspection after download.
The synthetic nature of the data may introduce artifacts not present in real-world advisory exchanges.
Provenance
Source
Soketlabs via Hugging Face.
Collection Method
Synthetically generated by Google Gemini 2.5 Flash.
Freshness
Last updated 2026-01-16 11:58:24; freshness should be verified.
License is unknown; users must verify permissions before use.