Browser Agent Evals

readme.md →

How today’s models perform on computer use benchmarks, compared on accuracy, cost, and speed.

15 runs, ranked by accuracy

accuracy per model

accuracy
project=stagehand-devbenchmark=hardbenchrows=15/57providers=4harnesses=7updated=2026-09-03T01:42:51Zrun=stagehand-evals
gpt-6-astracodex1st87.0%3rd320s$3.073
claude-fable-5-1claude code2nd84.0%552s$0.925
claude-opus-5claude code3rd78.0%678s$1.967
gpt-5.6-solcodex74.0%346s1st$0.502
gpt-5.6-soleve74.0%1st276s$0.856
claude-opus-5deep agents74.0%567s$1.502
claude-opus-5eve74.0%552s$1.809
gpt-5.6-solfx71.0%392s3rd$0.622
muse-spark-1.3mastra71.0%557s$2.085
gpt-5.6-soldeep agents71.0%349s1st$0.502
grok-4.6deep agents71.0%559s$2.955
claude-opus-5fx71.0%735s$10.023
gpt-5.6-solmastra71.0%2nd303s$1.412
grok-4.6mastra71.0%1154s$3.902
grok-4.6cursor71.0%668s$3.296

Run your own evals

same harness, your models, your benchmark

readme.md →

Want to get your model or harness evaluated?

talk to an engineer →

Your agent should be able to use a browser like you do.

The SDK for browser agents