Loading...
Loading...
Social network graphs, knowledge graphs, citation networks, molecular graphs for GNN, web link graphs
417 datasets
A set of benchmark tasks for node classification on knowledge graphs, designed to gauge progress in interpretable machine learning on relational and multimodal data. The datasets, created by Xander Wilcke of Vrije Universiteit Amsterdam, provide test and validation sets of at least 1000 instances, with some exceeding 10,000 instances. They are packaged in CSV format for easy consumption, with original RDF source data and pre-processing code provided for full provenance.
At least 1,000 test and validation instances per task, with some containing over 10,000 instances, provide a benchmark for evaluating machine learning models on knowledge graphs. The datasets, created by Peter Bloem of Vrije Universiteit Amsterdam, support both purely relational and multimodal learning tasks. They are packaged in CSV format for easy consumption, with original RDF source data and pre-processing code provided for full provenance.
22 online semi-structured interviews with men who have sex with men in Australia, conducted as part of a process evaluation for a social network-based HIV self-testing intervention. The data, authored by Ying Zhang and last updated in June 2026, includes thematic analysis of facilitators and barriers to implementation using the Consolidated Framework for Implementation Research (CFIR). Interviews explored perceptions of acceptability, feasibility, and contextual appropriateness of peer-led test kit distribution.
Food-web graph index data contains node and edge counts for arthropod species in maize fields. The dataset includes arthropod samples collected via pitfall traps, yellow traps, and plant assessments across different maize development stages in 2007-2008. It was authored by Zoltán Pálinkás and is available under an Open Access license.
Zachary P. Neal's dataset contains signed network backbones for the U.S. House and Senate across multiple sessions from 1973 to 2016. Relationships between congresspeople are inferred from bill co-sponsorship data using the Stochastic Degree Sequence Model. The data is structured as square matrices where cell values indicate positive (1), negative (-1), or no (0) relationships.
Simon Lee's study investigates the relationship between ResearchGate platform quality and scholars' research performance. The dataset, 60.9 KB in size, likely contains survey responses analyzing system, information, service, and collaboration quality. It was last updated on May 28, 2026.
50 randomly generated network files from a study on community detection algorithms. The dataset includes 30 ABCD graphs and 20 LFR graphs, created by Samin Aref and Mahdi Mostajabdaveh for their 2024 Journal of Computational Science article. Each network is provided in .gml format under a CC BY-NC-SA 4.0 license.
Great Britain's most accurate and authoritative path network dataset, produced by Ordnance Survey. It details footpaths through towns and cities, including information on responsible authorities and potential obstructions. The dataset was last updated on 2026-06-23.
Evolving Graph Attention Networks (EGAT) are proposed for learning from dynamic graphs where nodes and edges change over time. The 5.5 KB XLS file, authored by Yucai Jiang and last updated in June 2026, likely contains performance metrics from experiments comparing EGAT to other models. The results demonstrate that the proposed model outperforms state-of-the-art baselines on benchmark datasets.
Yucai Jiang published a dataset on figshare in 2026. The dataset contains statistics for benchmark datasets used to evaluate the EGAT model for dynamic graph representation learning. The file is 5.5 KB in size and is available in XLS format.
Helmholtz Knowledge Graph provides an RDF data dump from the Helmholtz Association's research infrastructure. The data is serialized in .ttl format and compressed with gzip, and is associated with major releases or data updates. It was authored by Volker Hofmann from Forschungszentrum Jülich and is available under an Open Access (green) license.
25 stakeholders were identified and analyzed for Abu Dhabi's National Bullying Prevention Strategy. The dataset, created by Alfan Al-Ketbi and last updated in May 2026, maps the influence and roles of 17 organizations using semantic and social network analysis. It benchmarks findings against international models from Finland, Norway, Australia, the USA, Japan, and Singapore.
A text mining-based biomedical knowledge graph constructed from all published literature related to human aging and longevity in PubMed. The dataset includes seven sets of files in JSON and CSV formats, such as literature metadata, entity and relation triples, and specific biomarker information. It was created by Zexu Wu and is available under an Open Access license.
IV-SAKG is a scenario-aware knowledge graph dataset constructed for vulnerability reasoning in industrial control systems. It contains entities, relations, and triples built upon public vulnerability sources like CVE, CISA, and CNVD. The dataset was authored by ling zhang and last updated on 2026-05-25.
VirtualMicrobes simulation data supports a Nature Communications Biology paper on contingent evolution. Jeroen Meijer from Utrecht University generated this data for a 2020 study. The data likely contains parameters and outcomes from simulations of evolving metabolic network topologies.
LetterSampo CKCC is a knowledge graph dataset of historical letters, documented on the Linked Data Finland page. The dataset is available via a public SPARQL endpoint hosted by the University of Helsinki. Author Eero Hyvönen is associated with the project.
The Wikidata Refugee Camps Dataset (WRCD) is a curated Linked Open Data resource derived from Wikidata. It provides structured information about refugee camps worldwide through a reproducible workflow based on archived SPARQL queries and documented curation procedures. WRCD was authored by Maher Asaad Baker and hosted on Harvard Dataverse, with a last recorded update in July 2026.
Data supporting a hierarchical framework for predicting acute, intra-day public transport demand surges during mega-events. The dataset was created by Juan Angel Lucio Rojas and relates to FIFA World Cup 2026 host cities Mexico City and the New York/New Jersey metropolitan area. It was last updated on May 27, 2026.
Legacy deposit note: This Figshare record is an earlier archive copy and is not the current recommended version of the Samuel & Audrey Photography Metadata Archive. The 39.8 MB ZIP file contains an earlier version of metadata records, CSV and JSONL exports, and documentation from a travel photography archive hosted on SmugMug. Author Samuel Jeffery deposited this legacy version, which is retained for continuity and historical version tracking.
Samuel Jeffery deposited a 23.2 KB legacy archive of a dataset directory for the Samuel & Audrey Media Network on figshare. The ZIP file contains a small machine-readable registry and README file from an earlier version of the network's dataset-index work. This record was last updated on 2026-05-30 and is retained for historical continuity, not as the current active directory.