Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
Aegis-Safety-DPO is a manually-curated preference dataset designed for Direct Preference Optimization and Group Relative Policy Optimization. Created by PolarAI, the dataset focuses on training models to refuse malicious requests rather than provide preachy or evasive responses. The dataset was last updated on March 1, 2026.
License is unknown; terms of use must be verified before application.