Project · benchmark
web-agent-comparison
A reproducible benchmark of 7 browser-automation MCP servers (Playwright, Chrome DevTools, browser-use, Firecrawl, Lightpanda, and others) run against frozen fixtures, so results replicate instead of drifting with the live web.
- Surfaced real vendor bugs: 30-second hangs, phantom listeners, TLS-fingerprint leaks.
- Frozen loopback fixtures -- every server gets the identical page, every run.
- A real-Chrome TLS fingerprint baseline, so "undetectable" claims get checked against an actual browser.
- Cross-platform evidence added in v1.0.2; v1.1 expands the fixtures to general web tasks.
Why it exists: my pipeline depends on these tools daily, and vendor READMEs are marketing. I measure the tools I depend on.
Where to look, in the repo: