Reflex vs Playwright MCP · Benchmark highlights
Up to 2.2x faster. Up to 86% fewer tool-reply tokens. Up to 75% fewer MCP requests.
Selected results from separate benchmarks versus Playwright MCP. Speed and estimated tool-reply tokens: July 2026 checkout. MCP requests: October 2026 profile task with an unreleased Reflex build. Results vary by workload.
July checkout: 62.7s to 28.1s and about 4,586 to 650 estimated tool-reply tokens. October profile: 4 MCP requests to 1. The token highlight measures tool replies, not total model usage. See the July results and the October comparison below.
Less waiting. Fewer tokens. Fewer requests.
Up to 1.7x faster and 44% fewer total model tokens versus Playwright MCP, based on the median profile-task results. The new Reflex build also completed that task with 75% fewer MCP requests. This is the unreleased local build compared with Reflex 0.12.2 and Playwright MCP 0.0.83.
Up to 1.7x
faster task completion
Profile: 40.0s to 23.5s
Up to 44%
fewer model tokens
Profile: 136,030 to 76,561
Up to 75%
fewer MCP requests
Profile: 4 requests to 1
The gains hold across the full measured workload
Across all three tasks, the new build averaged 20% less time, 21% fewer total model tokens, and 56% fewer MCP requests than Playwright MCP. Compared with the previous Reflex build, it averaged 16% less time, 11% fewer tokens, and 16% fewer requests.
| Tool | Time | Model tokens | MCP requests | Success |
|---|---|---|---|---|
| Playwright MCP | 40.9s | 167,568 | 6.89 | 9/9 |
| Reflex previous | 38.9s | 148,450 | 3.56 | 9/9 |
| Reflex new build | 32.8s | 132,325 | 3.00 | 9/9 |
Every task, side by side
| Task | Playwright MCP | Reflex previous | Reflex new build |
|---|---|---|---|
| Fill a profile | 40.0s / 136,030 / 4 | 32.8s / 103,326 / 2 | 23.5s / 76,561 / 1 |
| Place an order | 52.5s / 258,275 / 14 | 60.1s / 278,348 / 8 | 56.1s / 241,204 / 7 |
| Find a result | 27.2s / 106,758 / 3 | 19.1s / 76,314 / 1 | 22.6s / 79,078 / 1 |
The gains vary by task: Playwright MCP finished checkout faster than the new build, and the previous Reflex build was faster and used fewer tokens on lookup. These are measured results on three local tasks, not a promise for every website or workload.
How we measured
All tools used gpt-6-astra with medium reasoning in Codex, with fresh sessions and browser profiles. Tool order rotated between repetitions. The model chose its own actions; Playwright MCP could use its built-in batching. Independent page-state checks verified every run. Failed tool attempts and recovery work remain included.
Time includes model work, startup, tool calls, and teardown. Tokens are actual reported input plus output, including repeated context and cached input. They are not a dollar-cost measurement. MCP requests include tool calls and returned-file reads, not model reasoning cycles. These agent results are separate from the warm scripted tool timings below.
October 5, 2026: scripted tool comparison
The local changes improve Reflex's action batches: filling six fields fell from 1068 ms to 395 ms, and checkout fell from 1516 ms to 914 ms. The smaller reads and dependent API task changed little with the same options. Tuned Playwright MCP was faster than the new Reflex build on all five tasks, including 55 ms for the form and 227 ms for checkout. This run does not support a general claim that Reflex is the fastest tool.
Values below are median milliseconds for the whole flow, including navigation, across 7 measured repetitions per task and mode. All 245/245 runs verified the requested answer or final state across 5 tasks and 7 modes. The table shows five modes; the full run also included Reflex using the new output options and Playwright code batches with the default settle wait.
| Task | Reflex before | Reflex local changes | Playwright MCP standard | Playwright MCP tuned | Direct Playwright |
|---|---|---|---|---|---|
| Fill six fields | 1068 | 395 | 582 | 55 | 45 |
| Checkout | 1516 | 914 | 3245 | 227 | 209 |
| Extract a table answer | 351 | 341 | 587 | 44 | 39 |
| Read one precise table cell | 327 | 335 | 567 | 43 | 26 |
| Dependent API requests | 677 | 668 | 1612 | 55 | 46 |
Reflex before is the baseline at commit 12e7d38422dd. The local column is 0.12.2 + local performance changes, an unreleased build. The published CLI is still 0.12.2. Playwright MCP is version 0.0.83. Its standard mode uses multi-field browser_fill_form calls. The tuned mode uses browser_run_code_unsafe to batch code, snapshot-mode none, and timeout-settle 0. Direct Playwright is a browser scripting baseline with no MCP replies or model round trips.
New options can reduce replies and calls within Reflex. Capturing and reusing an API result reduced the dependent API workflow from 3 to 2 calls. Projecting the table answer reduced that workflow's returned text from 18,301 to 272 characters. This compares two ways to ask for the answer, not a mandatory old cost: old Reflex returned the same answer in 213 characters when given a precise cell selector. Current Playwright MCP action replies already use short summaries and file links. This run does not show a general token advantage over Playwright.
No language model was used. Selectors were written in advance for every tool. Runs used isolated headless profiles in the same installed Chrome 154.0.8037.93, a 1280x900 viewport, and an Apple M5 on macOS. Reflex and the direct baseline used Playwright 1.62.0; MCP used Playwright 1.64.0-alpha-1790635538000. Standard MCP kept its default settle wait; tuned MCP disabled that wait while still verifying each task's result. One warmup per task and mode was discarded, and mode order rotated between repetitions. Timings exclude process startup. These five local fixtures do not measure agent completion time or prove results on arbitrary websites. Required file reads count their actual text and one extra interaction. Token estimates use reply characters divided by four, excluding tool schemas and model input.
July 22, 2026: historical agent results
Historical setup: Reflex 0.7.2 versus Playwright MCP 0.0.78, one agent run per side per task. The July checkout source records a stale cart carried into the Playwright browser session. This was not a clean-profile, repeated comparison of current versions. Tool-reply tokens were estimated from characters divided by four.
This separate, historical run used a real agent (Claude Fable 5) to drive both tools through 6 fresh live-site tasks, start to finish, with a fresh agent per flow so neither side reused context. Reflex completed them in 2.8x fewer round trips, on 2.6x less context in aggregate (and 4-29x smaller on the page reads themselves), and finished 1.7x faster end to end in aggregate, 1.4x to 2.2x per flow, with all 6 tasks succeeding on both tools. The heaviest task in the set, a ten step checkout, took 2 calls against 11 and 7.1x fewer estimated reply tokens in that run.
In this July 22 run, Reflex finished 1.4-2.2x faster per flow. Fewer round trips coincided with lower end-to-end time. These single live-site runs do not establish a general speed advantage.
Fewer round trips coincided with lower end-to-end time in this historical run. These were single live-site runs measured 2026-07-22, so treat wall clock as indicative. They do not establish how the gap changes with page size or the number of actions.
Prefer to watch it happen? The playground replays three of those flows call by call, with the round trips and the tokens counting up on both sides. Nothing to install.
June 12, 2026: historical scripted comparison
Across the 3 flows verified head-to-head against Playwright MCP, the free default for Claude Desktop and other pure MCP clients, Reflex completed the same tasks in 2 tool calls per flow, 6 in total, where Playwright MCP needed 14. This was a scripted run with no model, so these are tool calls, not observed model turns.
June 12 verified flows used 2 Reflex calls each versus 3-6 Playwright MCP calls each, 6 versus 14 in total. No model ran.
June 12 historical reply-token estimates
Over the 3 verified June 12 flows, Reflex returned about ~13K estimated tokens of reply text where Playwright MCP returned ~56K: roughly 4x fewer. Estimates are reply characters divided by four, excluding tool inputs, schemas, and model reasoning. The benchmark explicitly requested full snapshots where its scripted loop needed a page read. These totals do not describe current MCP's automatic action replies, which can be short summaries and file links.
Historical reply-text estimates using characters divided by four. The run had no model.
In that historical measurement, one explicit browser_snapshot look at the W3C CSS Grid spec returned roughly ~169K estimated tokens. The condensed Reflex read returned roughly ~7K. These compare full page-reader outputs, not the text automatically sent after each action.
June 12 historical page-read sizes
On these four pages measured June 12, Reflex returned the smallest page-reader output. Playwright MCP's explicitly requested browser_snapshot returned a full accessibility tree; Reflex returned condensed text. Values are estimated tokens, calculated as characters divided by four. Bars are scaled within each chart. This is a historical reader comparison, not a claim about current automatic action replies or every website.
June 12 historical estimate: about 26x less snapshot text. Full explicit page reads, not automatic action replies.
June 12 reader estimates across the wider field
The same historical pages were also read with Playwright CLI and agent-browser. This table counts the full returned snapshot or file text using the same characters-divided-by-four estimate. A shell-capable agent can read selected file lines instead of all the text below, so these figures are not the minimum context each tool needs for a task.
| Page | PW MCP snapshot | PW CLI yml file | agent-browser -i | Reflex |
|---|---|---|---|---|
| Wikipedia article | ~31K tok | ~31K tok | ~4K tok | ~4K tok |
| GitHub repo | ~21K tok | ~21K tok | ~6K tok | ~3K tok |
| Hacker News | ~15K tok | ~14K tok | ~3K tok | ~2K tok |
| W3C CSS Grid spec | ~169K tok | ~171K tok | ~37K tok | ~7K tok |
June 12 flow results and Reflex features
- Verification is built into the batch. 2 calls per verified flow with expectations checked inside the guarded call, so a flow either lands or fails loudly without an extra round trip to confirm.
- Forensic failures. On the one shared miss, Reflex returned near-miss candidates, the console feed, and a screenshot. That gives an agent more information for a repair; this scripted benchmark did not measure repair turns.
- Tied-best first-try success (5/6), the highest in the field.
- Model-free saved-flow replay with self-healing. A saved flow can run without a model generating each step. Playwright code batches and direct scripts can also run without a model.
June 12, 2026 methodology
Measured 2026-06-12 on macOS, same machine, live public sites, with no LLM on either side. Versions pinned: @playwright/mcp 0.0.76, @playwright/cli 0.1.14, agent-browser 0.27.2 (Vercel). Same 6 real-world flows, same end-condition verification, each tool driven per its own agent documentation. Tool-side wall clock is a second or two slower per flow than the fastest tool, while the agent round trips that dominate real time are far fewer. Reproduce it yourself: node bench/headtohead.js (competitor tools install under bench/tools/).
See the install docs, Reflex vs Playwright MCP, or the pricing to get started.
Ready when you are
Up to 2.2x faster. Up to 86% fewer tool-reply tokens. Up to 75% fewer MCP requests.
Selected results from separate benchmarks versus Playwright MCP. Speed and estimated tool-reply tokens: July 2026 checkout. MCP requests: October 2026 profile task with an unreleased Reflex build. Results vary by workload. See the measurements.
Installed in 2 minutes. Your pages stay yours. Or install first and skip the account.