Hundreds of thousands of user-submitted prompts captured during the Mosscap prompt injection game at DEF CON 31. The collection includes raw text inputs ranging from successful adversarial attacks to general conversational queries directed at the AI character across multiple challenge levels.
Use Cases
- Train a binary classifier to detect prompt injection attempts using the raw prompt text
- Analyze the progression of attack complexity across the different Mosscap levels
- Fine-tune safety guardrails by testing model responses against the diverse set of adversarial inputs
Strengths
- Hundreds of thousands of raw text prompts submitted by users during a live security competition
- Data sourced directly from the Mosscap variant of the Gandalf game hosted at DEF CON 31
- Includes every prompt received by the system, encompassing both successful injections and non-adversarial noise