Loading...
Loading...
Drug-target interaction, molecular screening, ADMET, compound databases, pharmaceutical data
627 datasets
OpenML hosts a dataset for toxicity prediction, likely containing molecular descriptors or chemical structures. The dataset is tagged for cheminformatics and drug safety applications. Specific details on size, author, and update date are not provided.
OpenML hosts a dataset for toxicity prediction, likely containing molecular descriptors or chemical structures. The dataset is tagged for cheminformatics and drug safety applications. Specific details on size, author, and update date are not provided.
UCI's Toxicity dataset contains chemical compounds and their associated toxicity labels for predictive modeling. The dataset's size, specific features, and creation date are not specified in the available metadata. It originates from the UCI Machine Learning Repository, a known source for benchmark datasets.
A database compiled by Gaรฑรกn Aceituno, Judith, harvested from e-cienciaDatos and last updated in October 2025. It contains information on alkaloids, including their chemical groups, biological distribution, pharmacological activity, and adverse effects. A second sheet details nanomaterial modifiers used in electrochemical sensors for detecting these alkaloids.
Derify's dataset contains molecular structures for drug discovery, derived from the Druglike molecule datasets. The data was canonicalized using RDKit (2024.9.4) for structural consistency, and 33% of the dataset was randomly sampled and augmented using RDKit's Chem.MolToRandomSmilesVect function to enhance diversity. The dataset was last updated on September 9, 2025.
A dataset of 43 million canonicalized molecular structures derived from the Druglike molecule datasets. The dataset was created by Derify and last updated on September 9, 2025. To enhance diversity, 33% of the dataset was randomly sampled and augmented using RDKit's Chem.MolToRandomSmilesVect function.
SAIR provides 1,048,857 unique protein-ligand pairs and 5.2 million 3D structures curated from ChEMBL for drug discovery research. Created by SandboxAQ in collaboration with Nvidia and updated in August 2025, it pairs binding potency measurements with structural data.
IndoDiscourse is a multi-labeled Indonesian text dataset examining toxicity, polarization, and demographic information. The dataset was restructured and expanded by author Exqrch on October 31, 2024, and last updated on the Hugging Face platform in June 2025. It groups unique texts together and includes annotations from multiple annotators.
Pillbox contains metadata for oral solid dosage form medications, derived from FDA drug labeling. The dataset includes physical characteristics, active and inactive ingredients, National Drug Codes, and information about marketing firms. It was retired on January 28, 2021, and its final image library remains available for research.
DailyMed offers a standard resource of medication package inserts, known as Structured Product Labeling (SPL), from the U.S. Department of Health & Human Services. The repository is updated daily, with the most recent update in July 2025, providing current labeling information for drugs and supplements.
Drugbankrawparquet is a dataset published on the Hugging Face platform by user agenticx. The dataset was last updated on August 3, 2025. Its content likely contains raw data related to drugs and pharmacology, inferred from the title.
LiverTox provides current information on liver injury from prescription drugs, over-the-counter medications, and dietary supplements. The resource is maintained by the U.S. Department of Health & Human Services and was last updated in June 2025. It details the diagnosis, causes, frequency, and management of drug-induced liver injury.
An evaluation subset of the Jigsaw Toxic Comment Dataset containing Wikipedia talk page comments annotated for toxic behavior. The dataset is hosted by GuardrailsAI and was last updated on February 12, 2025. It is intended for model evaluation, with training recommended from the original Jigsaw dataset.
5,000 Ukrainian tweets were filtered and labeled via the Toloka.ai crowdsourcing platform, resulting in a balanced set of 2.5k toxic and 2.5k non-toxic texts. This dataset is part of the larger multilingual_toxicity_dataset collection. It was uploaded by ukr-detect and last updated in November 2024.
A red teaming dataset of human-crafted jailbreaking prompts for large language models. The dataset is distributed under the CC BY-SA 4.0 license by innodatalabs and was last updated on 2024-04-17. It is cited in a paper titled 'Benchmarking Llama2, Mistral, Gemma and GPT for Factuality, Toxicity, Bias and Propensity for...'.
3,300 human-annotated Thai tweets categorized into 2,027 toxic and 1,273 non-toxic samples. The corpus includes labels from three annotators guided by a 44-word dictionary and accounts for 506 tweets that are no longer publicly available via a TWEET_NOT_FOUND placeholder in the text field.
159,571 Wikipedia talk page comments labeled across six distinct categories of toxicity including toxic, severe_toxic, and identity_hate. Each record contains raw text from human discussions paired with binary indicators of offensive behavior as determined by human raters.
A 2021 research dataset from the WOAH 2021 workshop, created by Alexandros Xenos, John Pavlopoulos, and Ion Androutsopoulos. It is hosted on Hugging Face by the 'tasksource' account and was last updated in July 2023. The dataset is designed for studying how context influences the perception of toxicity in text.
Finnish Suomi24 comments annotated by human raters for toxic behavior. The dataset was created by TurkuNLP and was last updated on Hugging Face in June 2023.
RealToxicityPrompts contains 100,000 English sentence snippets extracted from the web by the Allen Institute for AI in 2020. It was developed to provide a standardized benchmark for researchers to quantify and mitigate the risk of neural toxic degeneration in large language models.