299 real-world robot demonstrations and synthetic vision-language data across two tasks: Cocktail and Open-World Visual Grounding. The collection includes reasoning annotations for each demonstration and is formatted for the LeRobot ecosystem to support the development of adaptive Vision-Language-Action models.
Use Cases
- Train a Vision-Language-Action (VLA) model using the reasoning annotations and robot trajectories in the cocktail folder
- Develop visual grounding models using the open-world visual grounding task data
- Fine-tune robotic policies using the demonstrations collected via the Universal Manipulation Interface (UMI)
Strengths
- 299 real-world demonstrations for the 'Cocktail' task collected using the Universal Manipulation Interface (UMI)
- Data is provided in the LeRobot format for compatibility with standardized robotic learning pipelines
- Includes reasoning annotations for each demonstration in the cocktail task folder to facilitate adaptive reasoning training