ACL 2026 benchmark dataset of 1,000 multimodal social activity scenarios created by adonaivera. It is designed to evaluate generative AI agents' ability to detect and correct unsafe behavior during iterative plan revision. Each scenario contains a natural-language social activity description and an unsafe hourly plan spanning 11 activities from 7 PM to 5 AM.
Use Cases
- Benchmarking AI agent safety detection capabilities based on multimodal social scenarios
- Evaluating iterative plan revision algorithms based on unsafe hourly plans
- Training AI agents to identify unsafe behavior in social activity descriptions
- Studying cross-modal consistency in generative agent social simulations
Strengths
- Contains 1,000 distinct multimodal scenarios for evaluation
- Each scenario includes an unsafe hourly plan with 11 specific activities
- Designed for a specific evaluation task presented at ACL 2026
Limitations
- Column-level documentation is absent; field semantics must be inferred after download
- Row count is unknown, which may limit suitability assessment
- Description metadata is limited; actual data quality requires manual inspection after download
Provenance
- Source
- huggingface
- Freshness
- Last updated 2026-06-29 19:10:48