GPT Capabilities for Extracting Tasks From Textual Process Descriptions
by Juergen Mangler / Technical University of Munich
Available on 1 platform
Sign in to view source links and access this dataset
Description
Three tables evaluate the capabilities of GPT 1, GPT 2, GPT 3, and GPT 3.5 for extracting tasks from textual process descriptions. Performance is measured using a range of similarity metrics, including semantic text similarity for extracted task sets and individual tasks. The dataset was created by Juergen Mangler of the Technical University of Munich and is derived from the Zenodo dataset with identifier 10.5281/zenodo.7783492.
Use Cases
Benchmarking LLM performance on task extraction based on the described similarity metrics.
Analyzing the effect of text paraphrasing on model output based on the mention of nine different paraphrasing methods.
Comparing contextual versus non-contextual semantic similarity scores for extracted tasks.
Studying the relationship between extracted task count and model capabilities.
Strengths
Evaluates four distinct GPT model generations (GPT 1, 2, 3, and 3.5).
Uses multiple evaluation metrics, including semantic text similarity for both sets and individual tasks.
Includes analysis of text variations through nine different paraphrasing methods per text.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count is unknown, which may limit suitability assessment.
Last update date is unknown; freshness unverified.
Provenance
Source
Technical University of Munich (Juergen Mangler)
Collection Method
Derived from evaluating LLMs on the Zenodo dataset (10.5281/zenodo.7783492).
License is listed as Open Access (green); specific terms should be verified from the source.