Loading...
Loading...
Social network graphs, knowledge graphs, citation networks, molecular graphs for GNN, web link graphs
415 datasets
CoDEx is a set of knowledge graph completion datasets extracted from Wikidata and Wikipedia, presented at EMNLP 2020. It comprises three knowledge graphs varying in size and structure, includes multilingual entity descriptions, and provides tens of thousands of hard negative triples. The benchmark was created by Tara Safavi of the University of Michigan to improve upon existing datasets in scope and difficulty.
A research paper from Columbia University describes experiences and practical lessons learned from conducting security and privacy measurements on a large-scale social network. The work provides recommendations on ethical planning, institutional review board processes, and data-gathering techniques for online social networks. The author is Iasonas Polakis.
From 14 to 28 July 2021, this dataset contains measurements from a network of 15 radiometers deployed during the LIAISE field campaign. It provides high-frequency solar spectral irradiance at 10 Hz, along with derived total shortwave irradiance and column-integrated water vapor at multiple temporal resolutions. All data is quality-controlled with metadata and flags, and Level 2 data is calibrated against high-quality references.
The dataset integrates three major manuscript databases into a single knowledge graph covering medieval and early modern periods. It unifies records from the Schoenberg Database of Manuscripts, the Bibale database, and the Medieval Manuscripts Catalogue from the Bodleian Libraries. This structured data is available via a public SPARQL endpoint and powers a dedicated semantic portal for research.
A citation network of 1,893 academic documents relevant to knowledge co-production, linked by 9,759 citation edges. The dataset was constructed by researchers at the University of Edinburgh using a systematic search of the Web of Science Core Collection. It includes publication years for all nodes and detailed metadata for a subset of 525 fully retrieved papers.
A dataset of 1842 molecules used to train a Graph Neural Network model for predicting PC-SAFT equation-of-state parameters from molecular structure. The model was developed by Wildson Bernardino de Brito Lima and last updated in May 2026, achieving a liquid density mean absolute percentage error of 8.21% on a test set of 1642 unseen molecules. It demonstrated improved accuracy over traditional group contribution methods for complex solvents like ionic liquids and deep eutectic solvents.
A Graph Neural Network model predicts Perturbed-Chain Statistical Associating Fluid Theory (PC-SAFT) parameters for 1842 molecules based on molecular structure. The model, trained on experimental vapor pressure and saturated liquid density data, demonstrated superior performance on a test set of 1642 unseen molecules and a subset of 122 molecules compared to a traditional group contribution method. Created by Wildson Bernardino de Brito Lima and last updated on 2026-05-28, the dataset supports the application of the PC-SAFT equation of state to complex, low-volatility solvents.
Experimental data from the HELIOS facility includes absolute transfer and transfer-induced fission cross sections for Uranium-238. The upload contains pre-sorted physics events, analysis code, an electronic logbook, and setup documentation. Samuel Bennett from the University of Manchester authored this dataset, which is shared under an Open Access license.
922 scientific articles on Social Network Analysis from South America, collected from the Web of Science database. The data was compiled by Adílson Luiz Pinto using bibliometric and scientometric methods to analyze publication and citation frequencies. The study covers 11 countries, with Brazil, Argentina, and Chile showing the highest centrality in the collaboration network.
High-resolution logbook data contains direct observations from 45 commercial fishing vessels collected in 2019. The dataset includes effort, fuel consumption, and annual landings by species, which were used to compute Fuel Use Intensity (FUI) and Carbon Footprint (CF) per vessel. It was created by Antonello Sala and combined with on-site energy audits to model daily fuel consumption based on vessel length.
GraphXAI is an open-source library providing a synthetic graph data generator and XAI-ready real-world datasets for evaluating Graph Neural Network explainers. The library includes ShapeGGen, which can generate benchmark datasets with varying graph sizes, degree distributions, and homophily levels, all accompanied by ground-truth explanations. It was created by Owen Queen of Harvard University Press and is available on paperswithcode.
A dataset of 1842 molecules used to train a Graph Neural Network model for predicting PC-SAFT equation-of-state parameters from molecular structure. The model was developed by Wildson Bernardino de Brito Lima and last updated in May 2026, achieving a liquid density mean absolute percentage error of 8.21% on a test set of 1642 unseen molecules. It demonstrated improved accuracy over traditional group contribution methods for complex solvents like ionic liquids and deep eutectic solvents.
Finland's Second World War is documented in a harmonized knowledge graph with over 14 million triples. The data covers the Winter War (1939-1940), Continuation War (1941-1944), and Lapland War (1944-1945). It is structured into subgraphs for events, actors, places, photographs, and other documentation, and powers the semantic portal WarSampo.
1,020 randomly generated network files in .gml format from a study on the Bayan community detection algorithm. The collection includes 500 ABCD graphs, 500 LFR graphs, 10 Erdos-Renyi graphs, and 10 Barabasi-Albert graphs. It was created by Samin Aref and colleagues for research published in 2022 and 2023.
A paper by James R. Clough of Imperial College London analyzes the Transitive Reduction of citation networks. The work uses this graph-theoretic operation as a tool to reveal causal structure and highlight edges more likely to correspond to information transfer. The associated paper is available via Open Access (green) license on arXiv.
Wastewater samples contain norovirus levels, ammonia, and orthophosphate concentrations. The data comprises 3,232 samples from 152 sewage treatment works and 220 samples from 11 sewerage network sites across England, collected between May 2021 and March 2022 for the Environmental Monitoring for Health Protection programme. It was aggregated by the Marine Environmental Data & Information Network.
18,440 catchments across the Tibetan Plateau are characterized by hydro-geomorphic unit hydrographs and 18 attributes. The dataset was produced by Yuhan Guo of Tsinghua University using the WFIUH extraction framework on the HydroBASINS dataset. It characterizes the rainfall and runoff response relationship.
22 online semi-structured interviews with men who have sex with men in Australia, conducted as part of a process evaluation for a social network-based HIV self-testing intervention. The data, authored by Ying Zhang and last updated in June 2026, includes thematic analysis of facilitators and barriers to implementation using the Consolidated Framework for Implementation Research (CFIR). Interviews explored perceptions of acceptability, feasibility, and contextual appropriateness of peer-led test kit distribution.
Food-web graph index data contains node and edge counts for arthropod species in maize fields. The dataset includes arthropod samples collected via pitfall traps, yellow traps, and plant assessments across different maize development stages in 2007-2008. It was authored by Zoltán Pálinkás and is available under an Open Access license.
Zachary P. Neal's dataset contains signed network backbones for the U.S. House and Senate across multiple sessions from 1973 to 2016. Relationships between congresspeople are inferred from bill co-sponsorship data using the Stochastic Degree Sequence Model. The data is structured as square matrices where cell values indicate positive (1), negative (-1), or no (0) relationships.