MC

microsoft/cuawright

Developer tools
6천 stars 품질 85 트렌드 85

A simple SWE style browser+desktop agent framework that achieves SOTA results on long horizon web tasks.

개요

Alongside browser automation, you can now run desktop tasks in an Ubuntu VM. Use cuawright web for browser tasks and cuawright desktop for desktop tasks. Existing webwright commands and Python imports still work. Turn coding models into browser and desktop agents - 📄 CUAWright: A Minimal Unified Interface for Digital Agents - 📝 Webwright: A Terminal Is All You Need For Web Agents - 🌐 microsoft.github.io/CUAWright CUAWright gives coding models a terminal to automate browsers and desktop applications. Use Webwright to browse the web and build reusable Playwright scripts, or run your own desktop task and reproduce OSWorld-V2 benchmarks in an Ubuntu VM. Start with the browser quick start or desktop quick start. Already got your favorite agents, and wonder how to make Claude Code, Codex, Hermes more capable in browser tasks? Consider adding CUAWright browser plugin/skill! - — renamed the project to bring browser and desktop agents together.

README

CUAWright

Webwright is now CUAWright. Alongside browser automation, you can now run desktop tasks in an Ubuntu VM. Use cuawright web for browser tasks and cuawright desktop for desktop tasks. Existing webwright commands and Python imports still work.

Turn coding models into browser and desktop agents

CUAWright gives coding models a terminal to automate browsers and desktop applications. Use Webwright to browse the web and build reusable Playwright scripts, or run your own desktop task and reproduce OSWorld-V2 benchmarks in an Ubuntu VM.

Start with the browser quick start or desktop quick start.

Already got your favorite agents, and wonder how to make Claude Code, Codex, Hermes more capable in browser tasks? Consider adding CUAWright browser plugin/skill!


📰 News

  • 2026-10-03 — Webwright → CUAWright: renamed the project to bring browser and desktop agents together. The Webwright browser runtime is now cuawright.webwright, and OSWorld desktop support lives at cuawright.desktop. Existing webwright commands and imports remain compatible. See desktop setup.
  • 2026-09-01 — Persistent step-by-step browsing and native run_command tool calls improve performance to 88.1% on Online-Mind2Web and 77.5% on Odysseys.
  • 2026-07-21 — Skill Factory: every solve leaves a script behind, distilled into reusable, verified, parameterized code skills that rerun standalone with no model (~40 s, zero tokens). On WebArena, reuse lifts held-out accuracy 55% → 70% (+15 pp). See the optional Skill Factory.
  • 2026-05-11 — Support Task2UI mode: Webwright completes the task and renders task results into an HTML-based web app you can easily view and reuse.
  • 2026-05-06 — Codex and Claude Code plugin manifests added; install via /plugin install cuawright@cuawright. Hermes Agent integration shipped; the same skills/cuawright-web/ folder now loads across Claude Code, Codex, and Hermes.
  • 2026-05-04 — Initial public release: ~1.5k LoC, OpenAI / Anthropic / OpenRouter backends, Playwright environment.




🎥 Demo

CUAWright demo

https://github.com/user-attachments/assets/2063be64-4d0a-40a2-9d1f-7906be43f2c9


📊 Performance

CUAWright beats the same model running in a different harness on every benchmark below, from desktop and CAD to long-horizon web tasks. See the paper for full details.

  • 🖥️ OSWorld-V2: 67.9% partial score with GPT-5.6 Sol (+5.2 over GPT-5.6 Sol alone) and 63.2% with GPT-5.5 (+15.7 over the official GPT-5.5 OSWorld agent).
  • 🌐 Web: 88.1% on Online-Mind2Web and 77.5% on Odysseys with GPT-5.4, against 83.4% and 33.5% for a GPT-5.4 GUI agent.
  • 📐 CAD: 79.0% Vision2Code mean IoU on BenchCAD with GPT-5.5, and 55.7% aggregate score on CADGenBench with GPT-5.6 Sol.
  • 💰 Cost: on OSWorld-V2, CUAWright raises the score while cutting API cost by $9.5 per task with GPT-5.6 Sol and $6.0 per task with GPT-5.5.

🗺️ Project Map

CUAWright/
├── pyproject.toml                      # package, dependency extras, CLI commands
├── setup.py                            # desktop provenance in wheels and source archives
├── src/
│   ├── cuawright/
│   │   ├── core/                       # shared API helpers
│   │   ├── webwright/                  # browser runtime
│   │   │   ├── agents/                 # browser agent loop
│   │   │   ├── environments/           # browser and terminal workspaces
│   │   │   ├── models/                 # OpenAI, Anthropic, OpenRouter backends
│   │   │   ├── tools/                  # browser sessions, images, self-reflection
│   │   │   ├── config/                 # stackable browser YAML configs
│   │   │   ├── run/                    # cuawright-web CLI and doctor
│   │   │   └── utils/                  # evidence, logging, serialization
│   │   └── desktop/                    # persistent desktop runtime
│   │       ├── agents/                 # desktop agent loop and call budget
│   │       ├── environments/osworld/   # VM commands, screenshots, and submission
│   │       ├── models/                 # standard Responses API transport and tools
│   │       ├── config/prompts.py       # desktop prompts
│   │       ├── setup.py                # pinned OSWorld setup and verification
│   │       ├── run/cli.py              # custom and official desktop workflows
│   │       ├── run/benchmarks/          # official task setup and evaluation
│   │       └── utils/                  # artifacts and provenance
│   └── webwright/__init__.py           # legacy Python import compatibility
├── extensions/skill-factory/           # optional cuawright-skill-factory distribution
│   ├── pyproject.toml                  # separate wheel; depends on the base package
│   └── src/cuawright/webwright/        # skill_factory/ and tools/skill_use.py
├── skills/cuawright-web/               # browser skill, commands, reference guides
├── .claude-plugin/                     # Claude Code plugin and marketplace manifests
├── .codex-plugin/                      # Codex plugin manifest
├── docs/                               # desktop setup and Skill Factory guides
├── tests/                              # browser, Skill Factory, compatibility tests
├── release/osworld/tests/              # desktop request, lifecycle, privacy tests
├── .github/workflows/                  # runtime and Skill Factory CI
├── assets/                             # showcase, trajectory viewer, figures, logos
├── LICENSE                             # MIT browser license
├── licenses/                           # Apache-2.0 desktop runtime license
└── NOTICE                              # imported runtime attribution

🚀 Browser Quick Start

Webwright lets a coding model launch browsers, inspect pages, and debug its work from a terminal. Script-based runs produce a reusable Python script. Live-browser mode keeps a page open across steps and returns an answer without creating a script.

Prerequisites

  • Python 3.10+
  • Chromium installed through Playwright
  • An API key for your chosen backend (OpenAI, Anthropic, or OpenRouter)

Install

pip install -e .
playwright install chromium

Script-based run

Export credentials for the configured backend (for example, OPENAI_API_KEY with model_openai.yaml or ANTHROPIC_API_KEY with model_claude.yaml). Then:

python -m cuawright.webwright.run.cli main \
    -c base.yaml -c model_openai.yaml \
    -t "Report the page title and the destination of its main link." \
    --start-url https://example.com \
    --task-id demo_openai \
    -o outputs/default

The image_qa and self_reflection tools use the run’s configured model. When running a tool separately, pass --model-config with an absolute path to a private YAML file containing your model: settings. Keep that file outside the checkout and results directory.

Live-browser mode

For incremental browsing with native OpenAI run_command tool calls, use:

cuawright-web main -c best_default_judge_json_persistent_cli.yaml -c model_openai.yaml \
  -t "" --start-url "" --task-id example -o outputs/example

This config keeps a Browserbase cloud session across commands. Set OPENAI_API_KEY, BROWSERBASE_API_KEY, and BROWSERBASE_PROJECT_ID before running. The agent runs shell commands, inspects the results, and checks its work before finishing. See the browser workflow guide for image reading and trajectory verification.

🚩 Flags

Flag Description
-c Config file(s) from src/cuawright/webwright/config/ (stackable).
-t Task instruction.
--start-url Initial page.
--task-id Output subfolder name.
-o Output directory.

🖥️ Desktop Quick Start

cuawright desktop runs tasks in an OSWorld-V2 Ubuntu VM. The agent uses a terminal to inspect applications, run commands, view screenshots, and save its work. You can give it your own task or run an official benchmark task with a score.

Requirements and installation

  • Python 3.12+ and a Responses-compatible model API key.
  • Docker and access to the gated official task and asset snapshots.
pip install -e ".[desktop]"

Before setup, accept access to the tasks and assets on Hugging Face, then log in:

pip install huggingface_hub
hf auth login
cuawright setup osworld --root ~/.cache/cuawright/osworld-v2-2026.08.08
cuawright verify osworld \
  --setup ~/.cache/cuawright/osworld-v2-2026.08.08/setup.json

Setup downloads the pinned source, task classes, assets, and VM by default; existing resources can be registered without copying. Put the API key in a user-owned file with mode 0600, outside source, task, asset, and result directories. Choose a new result directory whose parent already exists.

Run an arbitrary desktop task

cuawright desktop run \
  --setup /absolute/path/to/setup.json \
  --credentials /absolute/path/to/api-key \
  --results /absolute/path/to/results/custom \
  --model YOUR_MODEL \
  --instruction ""

This workflow has no evaluator or score. For an official scored task:

cuawright desktop reproduce osworld \
  --setup /absolute/path/to/setup.json \
  --credentials /absolute/path/to/api-key \
  --results /absolute/path/to/results/task-001 \
  --task-id 001 --model YOUR_MODEL --reasoning xhigh

For another Responses-compatible provider, pass its full API URL with --responses-url. See task selection and proxies for official task IDs and --proxy-config setup.

Each run saves its settings in manifest.json, the agent’s actions in trace.jsonl, and the outcome in result.json. Images viewed by the agent are embedded in the trace; desktop runs do not export separate screenshot files. Official benchmark runs also include the evaluator score. Files created in the VM are not automatically copied to the host; see custom tasks. See REPRODUCE.md for setup, smoke tests, and reproduction commands, or desktop setup and execution for more detail.


🔌 Use as a Plugin

CUAWright ships plugin manifests for both Claude Code (.claude-plugin/plugin.json) and OpenAI Codex (.codex-plugin/plugin.json), with the shared skill at skills/cuawright-web/ and slash commands at skills/cuawright-web/commands/. The host agent drives the Webwright loop natively — no extra LLM API key or cost beyond your host subscription. Hosts that read PNG screenshots natively skip the image_qa / self_reflection tools.

The bundled cuawright-web skill handles browser tasks. Desktop tasks use cuawright-desktop; see desktop setup.

Common runtime deps (install once after either path):

pip install -e .
playwright install chromium

Credits

Citation

If you use CUAWright or Webwright in your research or build on it, please cite:

@misc{lu2026cuawright,
  title         = {CUAWright: A Minimal Unified Interface for Digital Agents},
  author        = {Lu, Yadong and Lee, Theodore and Li, Yifei and Jang, Lawrence Keunho and Xue, Tianci and Su, Yu and Sun, Huan and Awadallah, Ahmed Hassan},
  year          = {2026},
  eprint        = {2610.04116},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  url           = {https://arxiv.org/abs/2610.04116}
}

@misc{webwright2026,
  title        = {Webwright: A terminal is all you need for web agents},
  author       = {Lu, Yadong and Xu, Lingrui and Huang, Chao and Awadallah, Ahmed},
  year         = {2026},
  howpublished = {\url{https://github.com/microsoft/Webwright}},
  note         = {GitHub repository}
}
View this README on GitHub

추천 도구

다른 키워드를 입력하거나 필터를 제거해 보세요.

설치

npx skillfish add microsoft/cuawright