Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
A 2019 dataset released with the ASE conference paper 'DIRE: A Neural Approach to Decompiled Identifier Naming'. It contains information from 3,195,962 functions decompiled from 164,632 unique binaries generated from C code scraped from GitHub. The dataset was created by Jeremy Lacomis of Carnegie Mellon University and is partitioned into 16 archives by binary hash.
Data is partitioned into 16 archives by the first hexadecimal digit of the binary's SHA-256 hash; users must download and combine archives as needed.