Sign in to view source links and access this dataset
Description
400,000 samples across five tasks were used to transfer a passive image editor into an autonomous, question-conditioned visual reasoning assistant. The dataset was created by BeichenZhang and last updated on 2026-05-25. It includes tasks such as Fine-grained Perception, Chart Understanding, Maze Solving, and Jigsaw Puzzle.
Use Cases
Fine-tuning vision-language models for fine-grained perception tasks based on the described task categories.
Training autonomous visual reasoning assistants based on the question-conditioned samples.
Benchmarking model performance on chart understanding and data interpretation tasks.
Developing agents capable of solving maze and puzzle navigation problems based on visual input.
Strengths
Contains 400,000 training samples, providing a substantial base for model fine-tuning.
Designed for a specific, complex task: transferring a passive image editor into an autonomous visual reasoning assistant.
Covers five distinct task types, including Fine-grained Perception and Chart Understanding.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count is unknown, which may limit suitability assessment.
Last updated 2026-05-25 11:45:27; freshness should be verified.
Provenance
Source
BeichenZhang on Hugging Face
Collection Method
Created as SFT (Supervised Fine-Tuning) training data for the ETCHR project.
Freshness
2026-05-25
License is unknown; users should verify terms of use before applying the data.