VINCIE-10M is a dataset for unlocking in-context image editing from video. It was created by researchers including Leigang Qu and Feng Cheng, using chain-of-thought prompting with a Vision-Language Model to annotate visual transitions between frames. The dataset was last updated on August 24, 2025.
Use Cases
- Training models for in-context image editing based on visual transition annotations.
- Benchmarking Vision-Language Models on video-to-image reasoning tasks.
- Developing systems that generate image edits based on contextual video frames.
Strengths
- Dataset construction uses a chain-of-thought prompting method with a VLM.
- Focuses on a specific task: unlocking in-context image editing from video.
Limitations
- Description metadata is limited; actual data quality requires manual inspection after download.
- Column-level documentation is absent; field semantics must be inferred after download.
Provenance
- Source
- huggingface
- Collection Method
- Created using chain-of-thought prompting with a Vision-Language Model to annotate visual transitions.
- Freshness
- Last updated 2025-08-24 14:10:41; freshness should be verified.