122,752 Query-Document pairs compiled from openly available academic datasets for training the VisRAG model. The dataset was created by openbmb and last updated on October 15, 2024. Data is organized in batches of 128, with all data within a batch sourced from the same original dataset.
Use Cases
- Training retrieval models for visual question answering based on the described Query-Document pairs.
- Benchmarking model performance on in-domain academic data from sources like ArXivQA and PlotQA.
- Fine-tuning multimodal models using the structured batch organization mentioned in the description.
Strengths
- Contains 122,752 Query-Document pairs, providing a substantial volume of training data.
- Data is sourced from multiple named academic datasets (e.g., ArXivQA, PlotQA, ChartQA).
- Data is organized with a batch size of 128, ensuring intra-batch dataset consistency.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count per constituent dataset is provided, but total file size and formats are unknown.
Provenance
- Source
- openbmb via Hugging Face.
- Collection Method
- Compiled from openly available academic datasets.
- Time Range
- null
- Freshness
- Last updated 2024-10-15 21:33:39.
- Geography
- null