Sign in to view source links and access this dataset
Description
271,000 DBpedia persons and organizations are aligned to their Twitter profiles via a pipeline of data acquisition, candidate acquisition, and candidate selection. The dataset, created by Yaroslav Nechaev, contains alignment data, entity data, and code in CSV and JSON formats from a release dated 15th May 2017. It bridges the Twitter social media world and the Linked Open Data cloud to enable knowledge transfer between Semantic Web and social media research.
Use Cases
Enriching Twitter user profiles with structured DBpedia knowledge based on entity alignments.
Training entity linking models for social media based on the candidate IDs and confidence scores.
Studying the representation of persons and organizations across social and structured data platforms based on the alignment data.
Strengths
Aligns 271,000 DBpedia entities to Twitter profiles, providing a substantial linkage.
Includes confidence scores for candidate matches, allowing for quality filtering.
Offers data in multiple formats (CSV, JSON, RDF) and provides pipeline code for recreation.
Limitations
Row count for the final aligned dataset is unknown, which may limit suitability assessment.
Last update date is unknown; freshness unverified.
Column-level documentation is absent; field semantics must be inferred after download.
Provenance
Source
Created by the SocialLink Pipeline, aligning entities from DBpedia and Twitter.
Collection Method
Alignment via data acquisition, candidate acquisition, and candidate selection phases.
Time Range
Dataset release dated 15th May 2017.
Data files are stored in compressed .gz format; associated .tql files can be accessed via MS SQL Server.