Sign in to view source links and access this dataset
Description
A text dataset of medical instructions and conversations processed through a multi-stage pipeline. The data was converted, deslopped, deduplicated using MinHash, filtered via rejection sampling, and grammar-corrected. It was created by ChaoticNeutrals and last updated on HuggingFace in November 2024.
Use Cases
Fine-tune a medical chatbot based on processed conversational data.
Benchmark instruction-following capabilities of language models in a medical context.
Train a model for medical question-answering based on the instruction-response pairs.
Study the impact of data cleaning techniques like deduplication and grammar correction on model performance.
Strengths
Data underwent a multi-stage processing pipeline including deduplication and grammar correction.
Dataset was last updated on 2024-11-13, suggesting recent maintenance.
Limitations
The specific number of rows, columns, and file formats are unknown, making scale assessment difficult.
Column-level documentation is absent; field semantics must be inferred after download.
The source and original time range of the medical conversations are not specified.
Provenance
Source
Processed from ShareGPT data, likely via The-Chaotic-Neutrals/ShareGPT-Formaxxing GitHub repository.
Collection Method
Converted, deslopped, MinHash deduplicated, rejection filtered, and grammar corrected.
Freshness
Last updated 2024-11-13 00:41:08.
License is unknown; users should verify terms of use before application.