Sign in to view source links and access this dataset
Description
Capybara is a multi-turn chat dataset in ShareGPT format curated by agentlans for use with LLaMA-Factory. Each line contains a conversation between a human user and a GPT AI, starting with a human message, and the dataset is exclusively in English. The dataset was last updated on July 19, 2024.
Use Cases
Fine-tuning conversational AI models based on multi-turn dialogue structure.
Training instruction-following models based on human-AI interaction examples.
Benchmarking model performance on tasks requiring context retention across conversation turns.
Studying dialogue patterns and response generation in a human-GPT chat format.
Strengths
Dataset is structured in the standardized ShareGPT format, facilitating integration with common training pipelines.
Conversations follow a defined pattern, starting with a human user message.
Language is specified as English only, providing a consistent linguistic domain.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count is unknown, which may limit suitability assessment.
Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
Source
agentlans on Hugging Face
Collection Method
Likely contains conversations between human users and GPT AI, but specific collection method is not detailed.
Freshness
Last updated 2024-07-19 01:33:11; freshness should be verified.
License is unknown; users must verify permissions before use.