Loading...
Loading...
Text classification, translation, QA, summarization, dialogue, sentiment analysis, language modeling, text corpora
49,952 datasets
Port Phillip Bay is a popular place for boating, fishing, yachting and other aquatic pastimes - but it is also a gateway for commercial shipping. This polyline dataset represents the areas for recreational boats to Steer Clear of shipping in Port Phillip. The dataset is published by the Victorian Ports Corporation under a CC-BY-4.0 license and was last updated on 2026-04-09.
Victoria, Australia, is the geographic scope for this dataset depicting zones of alluvial gold mineralisation and areas subject to historical alluvial mining. The data was collected by the Geological Survey of Victoria and is accompanied by other related geological datasets. It was last updated on 2026-04-09.
A theoretical physics dataset describing the discovery of a new subhydrogen atomic entity. It provides a self-consistent theoretical solution for its magnetic moment and decay mechanism, along with a general stability model. The dataset was authored by Sharp Liang and is archived on figshare under a CC-BY-4.0 license.
Monthly financial reports from Connecticut's online sports wagering operators, detailing revenue, taxes, and wager adjustments since operations began in October 2021. The data includes licensee-specific filings with corrections and amendments noted for accuracy. It is published by the State of Connecticut via its open data portal.
Bathythermograph (XBT) data from NOAA Ship David Starr Jordan captures ocean temperature-depth profiles in the North Pacific Ocean's coastal waters of California from April 25 to June 18, 1986. The data is processed to the NODC C116 format, recording temperature at non-uniform 'inflection point' depths to define the temperature curve, with standard instruments reaching 450 or 760 meters. This dataset provides a snapshot of ocean thermal structure for a specific cruise and time period.
Northwest Atlantic Ocean water column data were collected from April 21 to July 18, 1980 via moored current meters and bottle casts. The dataset, processed to the NODC F004 standard, likely contains measurements of salinity, pH, oxygen, nutrients, temperature, and current velocity components. It was submitted by the Atlantic Oceanographic and Meteorological Laboratory as part of the North East Monitoring Program (NEMP).
Santa Clara County's Medical Examiner-Coroner provides downloadable records for deaths from January 1, 2018 onward, covering both jurisdictional and reportable non-jurisdictional cases. The dataset includes details on cause, manner, demographics, and location for deaths occurring within the county. Records are updated nightly by the Santa Clara County government.
Current meter data from the High Energy Benthic Boundary Experiment (HEBBLE) project provides time series measurements of ocean currents in the North Atlantic Ocean. The dataset, processed to the NODC F015 standard, contains Eulerian current measurements from moored instruments deployed between March 19, 1983 and August 1, 1984. Data records likely include east-west and north-south current vector components, sensor depth, position, and may also report water temperature, pressure, and conductivity or salinity.
A land use plan for the municipality of Petershagen/Eggersdorf, specifically the 'Alte Gärtnerei/Hasenweg' area. The plan contains a land use concept for the entire municipal area, requiring implementation through development plans or by-laws. The data is provided by the Bundesamt für Kartographie und Geodäsie and is accessible via a WMS service.
Mining Remediation Authority data shows the location and attributes of licensed areas for underground, opencast, and underground coal gasification operations. The dataset is designed to be used in conjunction with other authority datasets for a complete view of mining activity. It was last updated on April 14, 2026.
Top 20 search terms and symptoms in online searches related to major depressive disorder by year (2022–2024). The dataset was created by Rikako Shimizu and last updated on May 12, 2026.
Top 5 online search terms related to major depressive disorder are listed for each year from 2022 to 2024. The dataset is published by Rikako Shimizu on figshare under a CC-BY-4.0 license and was last updated in May 2026. The data is provided in a small 9.5 KB XLS file.
Details the location and attributes of each known coal bearing subbasin or coal domain in Victoria. The data originates from the 'Victorian Coal - A 2006 Inventory of Resources' report. It is provided by the Department of Energy, Environment and Climate Action and was last updated in April 2026.
251 positive interarrival times in hours between unexpected failures, recorded between 10 February 2004 and 20 November 2008. The data comes from seven Enercon E-40 wind turbines located in the same wind farm in northern Germany and was authored by Mustafa Hilmi Pekalp.
25,000 high-quality synthetic examples designed for supervised fine-tuning of language models. The dataset was created by WithinUsAI and is formatted as JSONL with chat messages and metadata. It was last updated on May 18, 2026.
Ancient Chinese WordNet is a structured lexical database for vocabulary from the Pre-Qin period (before 221 BCE). It contains 38,781 word forms and 55,100 manually annotated senses, each linked to a synset in Princeton WordNet 1.6. The dataset was developed by Nanjing Normal University, with the project beginning in 2012.
2.0 MB of data supports a novel analytical formulation for predicting the stability and energy absorption of bi-level architected lattice metamaterials. The dataset, created by T. Mukhopadhyay and uploaded to figshare in April 2026, includes results from matrix-based calculations, finite element simulations, and experimental validations concerning curved beam architectures.
A 5.5 KB Excel file containing key results for BERT models with debiasing techniques. The dataset was authored by Sayyed Mohammad Pourya Momtaz Esfahani and last updated on May 4, 2026. It is shared under a CC-BY-4.0 license on the figshare platform.
sysevol-ai created a collection of natural-language code-search evaluation queries synthesized by LLMs. Each query is grounded on a real code symbol from a SWE-bench instance. The dataset is in active development, with counts representing a current snapshot.
Angular histograms of energetic neutral atom counts from the IBEX-Lo instrument. Each CSV file contains counts integrated over a full 8-day orbit for a single energy step and species category (hydrogen or oxygen/carbon). The data was produced by the National Aeronautics and Space Administration and was last updated on 2026-03 13.