AnyAudio-Judge Bench is a bilingual (English/Chinese) multi-domain benchmark for evaluating instruction-audio alignment, released with the paper "AnyAudio-Judge: A Dynamic Rubric-Based Benchmark and Evaluator for Audio Instruction Following". It contains 7,920 curated samples per language across 7 subsets. The dataset was created by author cucl2 and was last updated on June 2, 2026.
Use Cases
- Benchmarking audio instruction-following models based on the described rubric-based evaluation.
- Training evaluators for audio-language alignment using the provided positive and negative samples.
- Studying model failure modes in audio tasks based on the hard negatives generated via instruction swapping and attribute perturbation.
- Conducting cross-lingual (English/Chinese) audio-language model analysis based on the bilingual sample structure.
Strengths
- Contains 7,920 curated samples per language, providing a substantial evaluation corpus.
- Maintains a strict 1:1 positive-to-negative sample ratio per subset for balanced evaluation.
- Includes hard negatives created via instruction swapping and attribute perturbation, likely increasing benchmark difficulty.
- Each sample is annotated with a list of decomposed binary rubric items for detailed scoring.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment for large-scale training.
- Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
- Source
- cucl2 on Hugging Face, associated with the "AnyAudio-Judge" research paper.
- Collection Method
- Curated samples, with hard negatives generated via instruction swapping and attribute perturbation.
- Freshness
- Last updated 2026-06-02 08:03:32; freshness should be verified.