A collection of prompt injection attempts filtered from millions of user interactions with the Gandalf security game during July 2023. The data specifically isolates adversarial inputs designed to bypass system instructions, distinguishing them from general user queries through automated text analysis.
Use Cases
- Train detection models to identify prompt injection attempts by analyzing the text of adversarial samples
- Benchmark LLM guardrails against real-world 'ignore instructions' strategies identified in the dataset
- Perform linguistic analysis on the specific phrasing and structures used by humans to bypass system prompts in a gamified environment
Strengths
- Sourced from a pool of millions of user-submitted prompts collected in July 2023
- Focuses specifically on 'ignore instructions' style prompt injections used against the Gandalf LLM
- Filtered using OpenAI text analysis to isolate relevant adversarial samples from general user chatter