Loading...
Loading...
Drug-target interaction, molecular screening, ADMET, compound databases, pharmaceutical data
627 datasets
Real Toxicity Continuations is a text dataset for evaluating the toxicity of language model outputs. It was created by user 'sasha' and last updated on Hugging Face in July 2022. The dataset contains prompts and continuations, likely sourced from models like GPT-2, to measure the propensity of language models to generate harmful text.
Encompassing 200,000 text documents from The Pile, scored for toxicity using the Perspective API in May 2022. It is balanced with 100,000 of the most toxic documents and 100,000 randomly sampled documents.
Serving as for text classification, automatically processed by AutoTrain for the 'procell-expert' project. The data instances contain 'text' and 'target' fields, as shown in a sample describing antitumor activity research. The author is Mim, and it was last updated on April 29, 2022.
A curated subset of The Pile dataset focused on toxic text examples, balanced for training and evaluation. The dataset was created by researcher tomekkorbak and uploaded to Hugging Face in June 2022. It is part of a series of balanced subsets derived from the larger 825GB Pile corpus.
Featuring 200,000 text documents from The Pile, balanced for toxicity. It was created by selecting the 100,000 most toxic and 100,000 least toxic documents from a 7-million-document subset scored using the Perspective API. The dataset was authored by tomekkorbak and last updated in April 2022.
A collection of a scored subset of 2.2 million documents from The Pile, processed through the Perspective API on May 18-20, 2022. It was created by tomekkorbak and provides toxicity annotations for text chunks.
Designed for fine-tuning language models on protein-ligand binding affinity and contact prediction. It contains molecular data tagged with categories such as Molecules, SMILES, and Chemistry. The dataset was authored by jglaser and last updated in May 2022.
A filtered subset of the Pile dataset, focused on text with toxicity labels, curated by researcher tomekkorbak and hosted on Hugging Face. It contains approximately 100,000 text samples, as indicated by its size category, and was last updated in April 2022. The data is intended for training and evaluating language models on toxic content.
Toxicity Debug is a text dataset for evaluating language model safety, created by researcher tomekkorbak and hosted on Hugging Face. It was last updated in April 2022. The dataset's size is categorized as 'n1 K', indicating it contains over 1,000 entries.
Pile Toxicity Balanced2 is a text dataset designed for training and evaluating language models on toxic content. The dataset, created by researcher tomekkorbak, was uploaded to Hugging Face in April 2022. It is part of a series of datasets derived from The Pile, a large-scale text corpus used for AI development.
A dataset for fine-tuning language models on protein-ligand binding affinity prediction. It is associated with tags for molecules, SMILES strings, and chemistry, indicating a focus on molecular data. The dataset was last updated in March 2022.
Two categories of protein-ligand interaction data, binding affinity and contact prediction, are provided for language model fine-tuning. These labels enable the development of predictive models for biochemical interactions and structural contacts between proteins and ligands.
Known as titled 'Medication' and was authored by mrojas. It was last updated on June 7, 2021. The number of rows, columns, and specific data content are unknown.
The dataset contains chemical functional use data, associated ToxPrint descriptors, and EPI Suite properties, supporting a 2017 publication on high-throughput screening for functional substitutes. It includes 729 ToxPrint descriptors per chemical and model outputs such as confusion matrices and bioactivity indices. The data was compiled by the U.S. Environmental Protection Agency for research on chemical alternatives.
The Toxicity Reference Database (ToxRefDB) from the U.S. Environmental Protection Agency contains toxicity testing results for 474 chemicals, primarily pesticide active ingredients. It consolidates approximately 30 years and $2 billion worth of animal studies previously found only in paper documents.
2020 data from experiments comprising 110,674 rats presents neurotransmitter response patterns for 258 clinically approved and experimental neuropsychiatric drugs. The dataset, authored by Hamid R. Noori, was used to analyze links between molecular drug action and neurobehavioral effects, revealing mismatches between drug classifications and systems-level neurotransmitter patterns.
Calcium imaging results from probing ligand-binding properties of human and rat α7 nicotinic acetylcholine receptor (nAChR) mutants. The data was generated by transient co-expression of α7/α9 nAChR mutants with chaperones and the Case12 calcium sensor, followed by pharmacological analysis using fluorescence microscopy or a FLIPR reader. It includes determined affinities for acetylcholine and epibatidine for wild-type receptors and specific mutants at positions 117–119, 184, 185, 187, and 189.
GSK's Tres Cantos Antimycobacterial Set screening identified 50 drug-like compounds prioritized from 250,000 candidates. The dataset includes computational predictions of their mechanisms of action, generated by Maria Jose Rebollo-Lopez in 2020. It is intended to support open-source tuberculosis drug discovery.
Aggregating gene expression and high-content imaging data from primary human kidney cells exposed to 46 diverse toxicants. It was used to identify biomarkers for predicting nephrotoxicity and inferring mechanisms of toxicity via Random Forest machine learning and network analysis. The data includes mRNA levels of HMOX1 and SQSTM1, along with imaging features capturing cell morphology and nucleus texture changes.
A PhD thesis from 2018 presents statistical methods for analyzing ecotoxicological data to derive environmental quality guidelines. The research focuses on Antarctic toxicity data and proposes improvements for dose-response modeling and species sensitivity distribution (SSD) construction. The work was conducted by an author affiliated with the Australian Antarctic Data Centre (AU_AADC).