At a glance
V4 Agent — most accurate
V4 writes and runs code to complete long, multi-step browser tasks. It can research across many pages, compare options, follow detailed instructions, collect records, and save the results as a spreadsheet or in any format. Best for:- Bulk data collection
- Complex tasks that span many pages and sites
- Long, difficult instructions
- “Collect every document with its title, reference, and deadline.”
- “Create a spreadsheet of all the pricing plans across these five competitors.”
V3 Agent — fastest
V3 looks at the page and acts step-by-step like a human would. Use only when cost and speed matter more than accuracy. Best for:- Simple tasks, such as finding publicly available information
- Workloads where cost and speed are essential
- Open-ended research
- “Get a shipping quote for a 3 lb package between two zip codes.”
- “Check whether this product is in stock and what it costs right now.”
V2 Agent — legacy
Our first-generation agent remains available for compatibility. It is based on our open-source repository, but it is not actively maintained and is less accurate. Starting something new? Use V4 instead; it’s far more accurate.V4 completes the hardest tasks
V4 completed 76% of difficult, real-world browser tasks — nine points ahead of V3 and twenty-two ahead of V2.
Typical tasks cost $0.70–$1.64
Measured with the same model — Claude Opus 4.7 — across the last 30 days of production usage, the typical customer’s successful task costs $0.70 on V2, $0.87 on V3, and $1.64 on V4, taking about 1, 2, and 5 minutes respectively. V4 costs more per task because customers use it for more complex work, and V2’s tasks run shortest only because it gets the simplest ones. On identical hard tasks, V3 is the fastest agent (about 6 minutes, versus about 10 for V2 and V4).
Comparing accuracy, speed, and cost
