Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
85 million protein sequences were generated by clustering 400 million sequences from the Open Genomic dataset at 90% identity and 90% coverage using MMseqs2 linclust. The dataset, created by tattabio, provides a non-redundant protein sequence resource for bioinformatics research.
Column structure and available metadata are not described in the provided input; users should inspect the dataset page for full details.