UTexas Aptamer Database: 1,415 Sequences from 489 Papers (1990-2022)
by Ali Askari / The University of Texas at Austin
Available on 1 platform
Sign in to view source links and access this dataset
Description
A collection of 1,415 aptamer sequences extracted from 489 scientific papers published between 1990 and 2022. The dataset, a snapshot from the University of Texas Aptamer Database, includes multiple sequences per selection experiment, akin to a metagenomic approach. For each aptamer, information includes publication details, target, sequence, GC percentage, length, binding affinity, and application.
Use Cases
Predict aptamer binding affinity based on sequence features like GC percentage and length.
Analyze trends in aptamer target selection and applications over the described multi-decade time range.
Train sequence generation models for novel aptamers based on the provided nucleic acid composition and pool data.
Study the relationship between selection buffer conditions and reported binding affinity.
Strengths
Contains 1,415 aptamer sequences, providing a substantial corpus for analysis.
Spans 489 source papers published over more than three decades (1990-2022), offering historical perspective.
Includes multiple sequences per experiment, not just high-affinity binders, which likely increases sequence diversity.
Limitations
Row count is unknown, which may limit suitability assessment for large-scale modeling.
Column-level documentation is absent; field semantics must be inferred after download.
Last update date is unknown; freshness unverified beyond the August 2023 snapshot.
Provenance
Source
The University of Texas at Austin (UTexas Aptamer Database).
Collection Method
Extracted manually from the scientific literature.
Time Range
1990-2022
Freshness
Snapshot from August 2023; update frequency of the source database is unknown.
License is described as 'Open Access (green)'; specific terms should be verified from the source.