Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
A benchmark subset of 1,101 tables from the GitTables corpus curated for evaluating column type detection systems. It was created by Madelon Hulsebos of the University of Amsterdam for the SemTab 2021 challenge's CTA task. The dataset provides ground truth annotations linking table columns to semantic types from the DBpedia and Schema.org ontologies.
The download page for the full GitTables corpus is provided separately. Column names in the benchmark tables are anonymized.