Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
Frank F. Xu from Carnegie Mellon University created datasets for the paper "A Systematic Evaluation of Large Language Models of Code". The data includes test sets of approximately 100 files each in 12 programming languages, which were not included in The Pile training corpus. These test sets were used to evaluate models including Codex, GPT-J, GPT-Neo, GPT-NeoX-20B, CodeParrot, and the authors' PolyCoder model.
The primary data file is 'unseen_test_sets.tar.gz'; other files like '2-7B-150K.tar' are trained model checkpoints, not raw datasets.