kgbench: Node Classification Tasks on Knowledge Graphs
by Xander Wilcke / Vrije Universiteit Amsterdam
Available on 1 platform
Sign in to view source links and access this dataset
Description
A set of benchmark tasks for node classification on knowledge graphs, designed to gauge progress in interpretable machine learning on relational and multimodal data. The datasets, created by Xander Wilcke of Vrije Universiteit Amsterdam, provide test and validation sets of at least 1000 instances, with some exceeding 10,000 instances. They are packaged in CSV format for easy consumption, with original RDF source data and pre-processing code provided for full provenance.
Use Cases
Benchmarking relational graph models in isolation based on the purely relational task mode described.
Evaluating multimodal relational graph models based on the option to use multimodal information.
Training models for node classification tasks that require pooling information from several steps away in the graph.
Conducting link prediction experiments, as the datasets may also be used for this task.
Strengths
Test and validation sets contain at least 1000 labeled nodes, with some exceeding 10,000 instances, enabling precise performance measurement.
Includes both purely relational and multimodal task variants to evaluate different model types.
Provides full provenance with original RDF source data and pre-processing code, packaged in an easily consumable CSV format.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count and overall dataset size are unknown, which may limit suitability assessment.
Last update date is unknown; freshness unverified.
Provenance
Source
Vrije Universiteit Amsterdam
Collection Method
Benchmark tasks created from knowledge graph source data, processed and packaged for machine learning evaluation.
License is listed as Open Access (green); specific terms should be reviewed.