UltraMedical-Preference is a dataset from TsinghuaC3I that enhances the UltraMedical collection with preference annotations. It includes responses sampled from both open-source and proprietary models, annotated for user preferences. The dataset was last updated on August 20, 2024.
Use Cases
- Training reward models for reinforcement learning from human feedback (RLHF) based on annotated preference data.
- Benchmarking the performance of different LLMs on medical question-answering tasks based on the annotated model responses.
- Fine-tuning language models for improved alignment with user preferences in medical contexts based on the provided annotations.
Strengths
- Includes annotations for user preferences, which are a key component for alignment research.
- Leverages responses from a mix of models, including proprietary ones like gpt-3.5-turbo and gpt-4-turbo.
- Builds upon the established UltraMedical collection, suggesting a related context.
Limitations
- Description metadata is limited; actual data quality requires manual inspection after download.
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment.
Provenance
- Source
- TsinghuaC3I
- Collection Method
- Responses sampled from open-source and proprietary models, then annotated for preferences.
- Freshness
- Last updated 2024-08-20 13:16:46; freshness should be verified.