kgbench: amplus: Node Classification and Regression Benchmark Tasks
by Peter Bloem / Vrije Universiteit Amsterdam
Available on 1 platform
Sign in to view source links and access this dataset
Description
At least 1,000 test and validation instances per task, with some containing over 10,000 instances, provide a benchmark for evaluating machine learning models on knowledge graphs. The datasets, created by Peter Bloem of Vrije Universiteit Amsterdam, support both purely relational and multimodal learning tasks. They are packaged in CSV format for easy consumption, with original RDF source data and pre-processing code provided for full provenance.
Use Cases
Benchmarking relational graph models in isolation based on the purely relational task setting.
Evaluating multimodal relational graph models based on the inclusion of multimodal information.
Training models for node classification tasks based on the described node labeling focus.
Training models for node regression tasks based on the benchmark's stated scope.
Performing link prediction tasks based on the description that each dataset may also be used for this.
Strengths
Test and validation sets contain at least 1,000 instances, with some exceeding 10,000, enabling precise performance measurement.
Datasets include both purely relational and multimodal information, allowing for isolated and combined model evaluation.
Full provenance is provided with original RDF source data and pre-processing code.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count for the full dataset is unknown, which may limit suitability assessment.
Last update date is unknown; freshness unverified.
Provenance
Source
Vrije Universiteit Amsterdam
Collection Method
Benchmark tasks created for research, with original source data in RDF.
License is Open Access (green); specific terms should be reviewed.