Loading...
Loading...
Source code corpora, bug reports, vulnerability databases, network intrusion detection, malware samples
2,236 datasets
LibScan_DataSet is a 11.1 MB ZIP file containing the source code and data for the LibScan project. It was authored by Wang Yishun and last updated on April 28, 2026. The dataset is shared under a CC-BY-4.0 license on the figshare platform.
Replication Data for "Debt Maturity and Commitment to Firm Policies" contains the Matlab code used to solve models and generate figures and tables for the academic paper. The code was authored by Andrea Gamba and Alessio Saretto for their work forthcoming in the Review of Corporate Finance Studies. The dataset was last updated on June 1, 2026.
A briefing package prepared by the Office of the Privacy Commissioner of Canada for a Senate committee appearance on December 4, 2025. The document contains analysis and positions regarding Bill C-15, the budget implementation act. It was published on the Open Canada portal in April 2026.
Annual self-reported filings from property owners in New York City, capturing bedbug infestation history as required by Local Law 69 of 2017. The data includes counts of infested and eradicated units, building identifiers, and precise geographic coordinates. It is published by data.cityofnewyork.us and was last updated on 2026-03-15.
210,000 total samples, including 200,000 benign and 10,000 phishing instances, are available for security model training. The dataset is hosted on Kaggle, but its author, creation method, and specific features are not detailed. Column-level documentation is absent, requiring users to infer the data structure after acquisition.
A legacy report from the Bureau of Mineral Resources (BMR) committee outlining a forward marine program. The document is published by Geoscience Australia on data.gov.au and was last updated on 2026-05-14. No abstract or sample data is available for content verification.
xAFS is an evaluation dataset for agentic retrieval over realistic, cross-context personal file systems. Each data point is a synthetic-but-realistic person with a folder containing emails, Slack exports, meeting notes, lab notebooks, contracts, photos-described-as-text, journals, and code reviews. The dataset was created by supermemory and last updated on Hugging Face in May 2026.
McNdroid is a large-scale, longitudinal, multimodal dataset for Android malware detection designed to benchmark concept drift robustness. It spans samples collected from 2013 to 2025 and provides three complementary modalities: static feature vectors, API call graphs (GML), and JSON-based behavioral representations. The dataset was created by IQSeC-Lab.
3,000 curated training samples bridge command-line syntax and forensic evidence. The dataset is private and governed by a proprietary license, created by author dpevzner and last updated on May 5, 2026.
A 2.2 GB dataset created by G. Sai Chaitanya Kumar and last updated on 2026-04-21. It contains data for evaluating a lightweight intrusion detection system based on E-GhostNetV2-MobileNeXt, designed to secure the Internet of Health Things (IoHT). The study utilized three widely recognized public datasets for multiclass and binary classification analyses.
73,610 binaries spanning 248 open-source projects, compiled with multiple compilers and optimization levels for Linux and Windows. The dataset includes multi-year version histories and is linked to 329 CVEs. It was uploaded by author changliu8541 to Hugging Face, with a fix noted in May 2026.
A tabular dataset curated for evaluating predictive models on independent and identically distributed data, with the intended task of classification. The data originates from a 2014 study on phishing detection using associative classification data mining and is licensed under CC BY 4.0.
One official declaration from the 2018 G7 Charlevoix Summit outlines shared principles for tackling marine plastic pollution. The document was issued by the leaders of the Group of Seven and is archived by Global Affairs Canada. It is preserved for research and recordkeeping purposes only.
Record for source data hosted in the National Spectral Database (NSD) Aquatic Library. The dataset is part of a 2007 technical report for the Adelaide Coastal Waters Study, prepared by David Blackburn Environmental Pty Ltd and CSIRO Land and Water. It is managed by Geoscience Australia Data and was last updated in April 2026.
A spectral library for aquatic substrates collected from Adelaide coastal waters in 2003. The data is hosted in the National Spectral Database and was created as part of a remote sensing study for the Adelaide Coastal Waters Study Steering Committee. The final technical report was published in July 2007 by David Blackburn Environmental Pty Ltd and CSIRO Land and Water.
Canadian briefing materials prepared for the Associate Deputy Minister of National Defence and the Commissioner of the Canadian Coast Guard. The package was created for an appearance before the Standing Committee on National Defence on December 11, 2025. The dataset was last updated on April 13, 2026, and is provided by National Defence.
Global Affairs Canada produced a briefing package for the Minister of Foreign Affairs's appearance at the Standing Committee on Procedure and House Affairs (PROC) on Foreign Election Interference. The dataset is an HTML document last updated on April 28, 2026. It is published under the OGL-CA-2.0 license.
Code Security Vulnerability Dataset is a curated collection of 175,419 code samples labeled with 31 vulnerability classes, including 30 Common Weakness Enumeration (CWE) types and a 'safe' category. Labels are mapped to OWASP Top 10 2021 categories. The dataset is split into training, validation, and test sets and includes samples from C, C++, Python, JavaScript, Java, PHP, and Go.
Tax Increment Financing (TIF) Investment Committee Decisions data contains all projects reviewed by Chicago's TIF Investment Committee from May 2019 through December 2024. It documents committee decisions on funding requests for public infrastructure and private development projects. The dataset is published by data.cityofchicago.org and includes records from the committee's active period before meetings were suspended in January 2025.
Static analysis metrics of Portable Executable (PE) headers for classification. The dataset is hosted on Kaggle, but its author, organization, and creation date are unknown. The number of rows and specific features are also unspecified.