Assembled from a curated subset of web-browsing tasks derived from the WebVoyager benchmark, specifically modified for temporal stability. It provides a collection of interaction scenarios verified to remain functional and accessible until December 20, 2025.
Use Cases
- Benchmark autonomous web agents using the validated task scenarios to ensure consistent environment responses
- Test the proxy-lite model's navigation capabilities against a stable set of live web targets
- Analyze agent failure modes by comparing performance on this 2025-valid subset against the original WebVoyager tasks
Strengths
- Modified subset of the original WebVoyager task suite
- Guaranteed validity for all included web tasks through December 20, 2025
- Serves as the evaluation baseline for the proxy-lite model