Sign in to view source links and access this dataset
Description
HumaniBench is a benchmark for evaluating large multimodal models using real-world, human-centric criteria. It consists of over 32,000 image-question pairs across seven tasks, including visual question answering, multilingual QA, and visual grounding. The dataset was created by the Vector Institute, with examples annotated using GPT-4o drafts and verified by experts.
Use Cases
Benchmarking model performance on open and closed visual question answering tasks based on the described image-question pairs.
Evaluating multilingual capabilities of vision-language models based on the multilingual QA task.
Testing model robustness and reasoning on adversarial or challenging examples as described in the task list.
Assessing model alignment with ethical guidelines through the described ethics evaluation component.
Strengths
Contains over 32,000 image-question pairs, providing a substantial evaluation corpus.
Covers seven distinct evaluation tasks, including multilingual QA and visual grounding.
Examples were annotated with GPT-4o drafts and then verified by experts, suggesting a quality control process.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count is unknown, which may limit suitability assessment.
Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
Source
Vector Institute
Collection Method
Examples annotated with GPT-4o drafts and verified by experts.
Freshness
Last updated 2026-04-29 18:45:18; freshness should be verified.
License is unknown; terms of use must be verified before application.