MW

microsoft/webwright

Developer tools
5.8K stars Quality 55 Trend 55

A simple SWE style browser agent framework that achieves SOTA results on long horizon web tasks.

Overview

Turn Your Coding Models to Be State-of-the-art Browser Agents - 📝 Webwright: A Terminal Is All You Need For Web Agents - 🌐 microsoft.github.io/Webwright Webwright gives LLM a terminal where it can launch multiple browser sessions to inspect the page and complete a web task. It captures and inspects page screenshots/states only when needed. It enforces each web task to be completed end-to-end within a re-runnable Python script, i.e. your web agent browsing history is a single code file. No multi-agent system, no graph engine, no plugin layer, no hidden orchestration — just a terminal, a browser, and a model. Already got your favorite agents, and wonder how to make Claude Code, Codex, Hermes, OpenClaw more capable in browser tasks? Consider adding Webwright plugin/skills! - — Support Task2UI mode: Webwright completes the task and renders task results into an HTML-based web app you can easily view and reuse.

README

Webwright

Turn Your Coding Models to Be State-of-the-art Browser Agents

Webwright gives LLM a terminal where it can launch multiple browser sessions to inspect the page and complete a web task. It captures and inspects page screenshots/states only when needed. It enforces each web task to be completed end-to-end within a re-runnable Python script, i.e. your web agent browsing history is a single code file. No multi-agent system, no graph engine, no plugin layer, no hidden orchestration — just a terminal, a browser, and a model.

Already got your favorite agents, and wonder how to make Claude Code, Codex, Hermes, OpenClaw more capable in browser tasks? Consider adding Webwright plugin/skills!


📰 News

  • 2026-05-11 — Support Task2UI mode: Webwright completes the task and renders task results into an HTML-based web app you can easily view and reuse.
  • 2026-05-06 — Codex and Claude Code plugin manifests added; install via /plugin install webwright@webwright. OpenClaw and Hermes Agent integrations shipped; the same skills/webwright/ folder now loads across Claude Code, Codex, OpenClaw, and Hermes.
  • 2026-05-04 — Initial public release: ~1.5k LoC, OpenAI / Anthropic / OpenRouter backends, Playwright environment.




🎥 Demo

https://github.com/user-attachments/assets/4ed94cd5-11be-4daa-b2d7-1260a803baca


📊 Performance

State-of-the-art on two real-website benchmarks with a 100-step budget — see the blog post for full details.

  • 🏆 Online-Mind2Web (300 tasks): 86.7% with GPT-5.4 — highest among open-sourced harnesses in the AutoEval category. Claude Opus 4.7 reaches 84.7%, and is stronger on the hard split (80.5% vs. 76.6% for GPT-5.4 at N=100).
  • 🚀 Odysseys (200 long-horizon tasks): 60.1% with GPT-5.4 (avg. 76.1 steps) — +15.6 points over the prior SOTA (Opus 4.6 at 44.5%, using vision based approach and persistent browser) and +26.6 points over base GPT-5.4 (33.5% using xy-coordinate prediction and persistent browser).
  • 🧠 Code-as-action beats coordinate prediction: Webwright substantially outperforms a reproduced GPT-5.4 screenshot+xy-coordinate baseline across all difficulty splits.
  • 🧰 Small models + reusable tools: generated scripts can be packaged as parameterized CLI tools — even Qwen-3.5-9B completes tasks well on Online-Mind2Web sites with 5+ tools available.

🗺️ Project Map

webwright/
├── pyproject.toml           # package: webwright
├── src/webwright/
│   ├── run/cli.py           # CLI entrypoint (`webwright`)
│   ├── agents/default.py    # core agent loop
│   ├── environments/        # Playwright browser workspace
│   ├── tools/               # image_qa, self_reflection
│   ├── models/              # openai_model, anthropic_model, base
│   ├── config/              # base.yaml, model_openai.yaml, model_claude.yaml
│   └── utils/
├── assets/
│   └── task_showcase/       # tiny Flask dashboard for repeatable runs
│       ├── app.py
│       ├── templates/       # dashboard.html, task.html
│       └── tasks// # task.json + report.json per task
├── tests/
└── outputs/                 # run artifacts (trajectories, screenshots)

📰 Task Showcase (repeatable runs as a dashboard)

A tiny Flask app under assets/task_showcase/ consolidates Webwright runs for repeatable odyssey tasks (deals, inventory, listings, job boards, weather, etc.) into a single dashboard. Each task ships only two files — task.json (metadata) and report.json (curated, structured output: sources + result sections like tables, lists, summaries) — and the templates render them generically, so adding a new task is just dropping a new folder in assets/task_showcase/tasks/.

pip install flask
python assets/task_showcase/app.py    # http://127.0.0.1:5005

To have Webwright produce a renderer-ready task folder at runtime, stack the Task Showcase overlay:

python -m webwright.run.cli \
    -c base.yaml -c model_openai.yaml -c task_showcase.yaml \
    -t "" \
    --task-id my_repeatable_task \
    -o outputs/default

Note: report.json is only generated when -c task_showcase.yaml is included. A plain base.yaml run produces trajectory.json and debug artifacts but no report.json.

The run writes task_showcase/tasks//task.json and report.json inside the output workspace. Render those generated files without copying them back into the repo:

python assets/task_showcase/app.py \
    --tasks-dir outputs/default//task_showcase/tasks

🚀 Quick Start

Prerequisites

  • Python 3.10+
  • Chromium installed through Playwright
  • An API key for your chosen backend (OpenAI, Anthropic, or OpenRouter)

Install

pip install -e .
playwright install chromium

Run

Export credentials for the configured backend (for example, OPENAI_API_KEY with model_openai.yaml or ANTHROPIC_API_KEY with model_claude.yaml). The image_qa and self_reflection tools use the same configured model by default, so an Anthropic run does not require an OpenAI key. Then:

python -m webwright.run.cli \
    -c base.yaml -c model_openai.yaml \
    -t "Search for flights from SEA to JFK on 2026-08-15 to 2026-08-20" \
    --start-url https://www.google.com/flights \
    --task-id demo_openai \
    -o outputs/default

🚩 Flags

Flag Description
-c Config file(s) from src/webwright/config/ (stackable).
-t Task instruction.
--start-url Initial page.
--task-id Output subfolder name.
-o Output directory.

🔌 Use as a Plugin

Webwright ships plugin manifests for both Claude Code (.claude-plugin/plugin.json) and OpenAI Codex (.codex-plugin/plugin.json), with the shared skill at skills/webwright/ and slash commands at skills/webwright/commands/. The host agent drives the Webwright loop natively — no extra LLM API key or cost beyond your host subscription. Hosts that read PNG screenshots natively skip the image_qa / self_reflection tools.

Common runtime deps (install once after either path):

pip install -e .
playwright install chromium

📃 Trajectory Comparison & Viewer

You can run the same tasks using the Webwright harness and its Codex / GitHub Copilot skill variant, and see how token usage and trajectories stack up between different harnesses. The trajectory viewer supports Codex, GitHub Copilot and Webwright harness traces.

How to use

cd assets/compare_trajectory/
python3 -m http.server

Open the webpage in your browser and upload the Webwright raw_responses.jsonl and attach trajectory.json to view. Then on the other side you can upload your Codex or GitHub Copilot trace.

Obtaining Codex traces:

ls ~/.codex/sessions/2026/MONTH/DAY/SESSION_ID.jsonl

Obtaining GitHub Copilot traces:

/export file session
-> session.md is the uploadable trace

Quick Comparison

“Find the cheapest used 8-cylinder bmw made between 2005-2015 and priced from 25,000 to 50,000 dollars with mileage less than 50,000 miles or less.”

Tokens Webwright Harness (Local Browser Mode) Codex Webwright Skill
Input 420,433 3,271,143
Output 3,593 20,040
Reasoning 0 4,410
Cached 217,216 3,081,3440
Total 424,026 3,291,183

Individual runs and results may vary.


Credits

Citation

If you use Webwright in your research or build on it, please cite this repository:

@misc{webwright2026,
  title        = {Webwright: A terminal is all you need for web agents},
  author       = {Lu, Yadong and Xu, Lingrui and Huang, Chao and Awadallah, Ahmed},
  year         = {2026},
  howpublished = {\url{https://github.com/microsoft/Webwright}},
  note         = {GitHub repository}
}
View this README on GitHub

Recommended Tools

Try a different keyword or remove a filter.

Install

npx skillfish add microsoft/webwright