Loading...
Loading...
Source code corpora, bug reports, vulnerability databases, network intrusion detection, malware samples
2,230 datasets
An Australian government report from the Bureau of Mineral Resources (BMR) committee outlining a forward marine program. The report is published via the Australian Ocean Data Network and was last updated on 2026-06-16. Its content is available in HTML and PDF formats.
A large-scale, quality-filtered cybersecurity corpus designed for continual pre-training of large language models. It contains approximately 15.5 billion tokens, including ~4.6B English text, ~5.5B Chinese text, ~4.7B code, and ~0.84B seed tokens. The dataset was created by WhitzardAgent and last updated on June 6, 2026.
Police and fire calls for service from the City of Cincinnati's computer-aided dispatch system. The dataset includes proactive and reactive incident data, updated every 15 minutes. Data is created and published by the City of Cincinnati's Office of Performance and Data Analytics.
A collection of code reviews mined from the GitHub and Gerrit platforms for analyzing the performance of Modern Code Review (MCR) techniques. The dataset was created by prahar pandya and is available under an Open Access (green) license. The specific size, time range, and detailed structure of the data are not provided in the available metadata.
A published dataset replicates a student experiment comparing exploratory testing and test-case-based testing methods. The dataset was created by Juha Itkonen and is shared under a Creative Commons Attribution-ShareAlike license. Temporal coverage and dataset size are not specified in the provided metadata.
Business license records for the City of Cincinnati, covering categories such as Amusement Games, Antique Dealers, and Massage Services. The dataset is processed and published daily by the city's Treasury Division and Office of Performance and Data Analytics. It includes administrative data with added geocoding and neighborhood information.
Vinaytosh Mishra published raw data supporting a manuscript titled 'Understanding the Influence Structure of Healthcare Cybersecurity Factors: A Fuzzy DEMATEL Approach' on figshare. The dataset is 27.3 KB in size and was last updated on May 30, 2026. It is available under a CC-BY-4.0 license.
Research by Dr. Adil Al Balushi measures the influence of career growth on turnover intention, mediated by employees' organizational commitment. The dataset is available via the paperswithcode platform under an Open Access (green) license. Specific details on the number of observations, features, and collection timeframe are not provided in the input metadata.
Measured data for antireflection coating design for a reverse engineering challenge. The dataset was authored by Jennifer D. Kruschwitz and is available under an Open Access license via paperswithcode.
A descriptive study authored by Glorin Sebastian investigates awareness, cybersecurity risks, and controls within the Metaverse. The work proposes a regulatory framework for this emerging digital space. It is available under an Open Access license on the paperswithcode platform.
Building permits applied for in Cincinnati since January 1, 2010, entered by the Buildings & Inspections department. The City of Cincinnati publishes this data with daily refresh frequency and provides an associated dashboard on CincyInsights. Data processing includes address verification, geocoding, and the addition of administrative areas like neighborhoods and police districts.
Raw data supporting the research paper 'Global convergence of dominance and neglect in flying insect diversity'. The dataset was contributed by author Amrita Srivathsan and is available under an Open Access (green) license. Specific details on the number of records, columns, and file formats are not provided in the metadata.
Anisha Islam provides data supporting a paper on the evolution of software testing practices. The dataset likely contains metrics or records extracted from Java project histories to analyze testing trends. Its specific size and temporal coverage are not detailed in the provided metadata.
Survey data from 502 visitors to Deosai National Park, Pakistan, exploring psychosocial factors influencing wildlife conservation intentions. The data was collected by Shakir Ali and analyzed using PLS-SEM to test an extended theory of planned behavior framework. The dataset was last updated on 2026-04 14.
Materials prepared for deputy heads for an appearance before the Standing Senate Committee on Foreign Affairs and International Trade (AEFA). The dataset is provided by Global Affairs Canada under the OGL-CA-2.0 license and was last updated on 2026-05-27. The content likely contains briefing notes, talking points, and background information for parliamentary committee hearings.
Cincinnati Building Permits Timeline tracks the progress of building and inspection workflows, including permits and zoning approvals. The dataset is maintained by the Cincinnati Area Geographic Information Systems (CAGIS) and individual City departments. It is refreshed daily.
Indicadores de Compromiso (IoCs) is a structured repository of cybersecurity threat indicators associated with Advanced Persistent Threat (APT) groups. The dataset includes indicators such as IP addresses, domains, URLs, and file hashes, collected from technical analysis, security event correlation, and threat intelligence sources. It is hosted on the DetecTIC platform and was last updated on 2026-04-14.
Properties registered in Cincinnati's vacant and foreclosed property program are tracked to ensure code compliance. The dataset is stored and maintained by Cincinnati Area Geographic Information Systems (CAGIS) and updated daily by the Department of Buildings & Inspections.
Fleet Services Division of Cincinnati's Public Services Department provides records of all vehicle maintenance and repair work orders from January 2008 onward. The dataset includes details on work types, costs, labor hours, and timestamps for service milestones. Data is refreshed daily by the Office of Performance and Data Analytics.
1.664 million cleaned and labeled source code samples across 16 programming languages, curated for language identification tasks. The dataset was created by author kaushik-harsh-99 and is hosted on Hugging Face. It was last updated on May 30, 2026.