MODUS is a large-scale dataset with pixel-aligned samples across 15 modalities. The dataset includes modalities covering appearance, geometry, structure, segmentation, detection, text, and learned features. It was created by epfl-vilab-modus and was last updated on July 5, 2026.
Use Cases
- Training any-to-any multimodal models based on the 15 aligned modalities.
- Developing cross-modal generation systems based on pixel-aligned appearance and geometry data.
- Training or evaluating segmentation models based on the included SAM and detection-based segmentation masks.
- Building vision-language models based on aligned RGB images and captions.
- Conducting research on modality fusion based on the aligned structure and feature representations.
Strengths
- Aligns 15 distinct modalities per sample, as stated in the description.
- Covers a broad range of data types including appearance, geometry, structure, segmentation, and detection.
Limitations
- Description metadata is limited; actual data quality requires manual inspection after download.
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment.
Provenance
- Source
- epfl-vilab-modus
- Freshness
- Last updated 2026-07-05 17:18:35; freshness should be verified.