Sign in to view source links and access this dataset
Description
36,115 synthetic tabular datasets containing over 72 million rows were generated for the TEXR project. Each dataset contains 2,000 rows and between 5 and 33 features, paired with JSON metadata describing its topic, features, and Bayesian-network structure. The collection was created by eddyliu-hf and last updated on 2026-06-26.
Use Cases
Benchmarking tabular data generation models based on the diverse topics and configurations described.
Testing machine learning algorithms on synthetic data with known Bayesian-network structures.
Studying the properties of synthetic datasets with varying feature schemas and value ranges.
Developing methods for data augmentation using controlled synthetic data generation.
Strengths
Large scale with 36,115 distinct datasets and over 72 million total rows.
Structured metadata for each dataset includes topic, feature descriptions, value ranges, and Bayesian-network structure.
Controlled generation with each dataset containing exactly 2,000 rows and between 5 and 33 features.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Data is entirely synthetic, which may limit direct applicability to real-world problems.
Last updated 2026-06-26 12:10:46; freshness should be verified.
Provenance
Source
huggingface
Collection Method
Synthetically generated for the TEXR project.
Freshness
2026-06-26 12:10:46
License is unknown; terms of use should be verified before application.