Sign in to view source links and access this dataset
Description
9,842 rows of instruction-response pairs for supervised fine-tuning, derived from two distinct Fable-5 corpora. The dataset was created by the author 'calmasacow' and was last updated on July 5, 2026. It combines 4,659 rows with added reasoning blocks and 5,183 rows of pure tool-use examples, verified to have no overlap in user-side content.
Use Cases
Fine-tuning language models for chain-of-thought reasoning based on the 'agentic-distill-fable-5-sft' corpus with added reasoning blocks.
Training models on tool-use interactions based on the 'fable-tool-use-sft' corpus derived from Claude Code sessions.
Creating a combined training set for multi-task instruction-following models by merging two distinct SFT corpora.
Studying the impact of explicit reasoning traces versus pure tool-use demonstrations on model performance.
Strengths
Contains 9,842 total rows, a union of two distinct corpora.
Comprises 4,659 rows with post-hoc reasoning blocks and 5,183 rows of pure tool-use examples.
Verified to have 0% overlap between the two source corpora via SHA-256 hashing of user-side content.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count is known, but other metadata like file formats, size, and license are unknown.
Data may reflect source bias inherent to the specific Fable-5 and Claude Code sessions used.
Provenance
Source
Combination of 'lordx64/agentic-distill-fable-5-sft' and 'lordx64/fable-tool-use-sft' from Hugging Face.
Collection Method
Union of two corpora, deduplicated by user-side content. Reasoning blocks were added post-hoc to one corpus by Glint-Research.
Freshness
Last updated 2026-07-05 18:19:55; freshness should be verified.
License is unknown, which may restrict commercial use or redistribution.