Sign in to view source links and access this dataset
Description
An LLM policy drove an agent browser over 642 tasks across 15 real websites to collect transitions for training world models. Each transition consists of a before screenshot, an action, and an after screenshot, capturing the consequence of web interactions. The dataset was created by sudac and last updated on June 8, 2026.
Use Cases
Train a world model to predict webpage changes based on the screenshot-action pairs described.
Evaluate the accuracy of a web action simulator using the before-and-after screenshot transitions.
Benchmark multimodal AI agents on tasks derived from the WebVoyager task set mentioned.
Fine-tune models for web navigation by learning from real browser interaction sequences.
Strengths
Data was collected from 642 tasks across 15 real websites, providing a diverse set of web interactions.
Transitions include multimodal data (screenshots and actions) specifically designed for world model training.
The collection method used an advanced LLM policy (gpt-5.4-mini) driving a browser agent, suggesting automated, scalable generation.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count is unknown, which may limit suitability assessment.
Freshness should be verified as the last update was in 2026.
Provenance
Source
huggingface
Collection Method
An LLM policy (gpt-5.4-mini) drove the vercel-labs/agent-browser over the WebVoyager task set.
Freshness
Last updated 2026-06-08 20:38:05
License is unknown; terms of use must be verified before application.