AM

antiv/mate

Developer tools
87 stars 0 forks 品質 35 トレンド 35

You built an agent. Now you need to tune the prompt. Swap the model. Restrict access for specific users. Figure out what it's actually costing you.

概要

You built an agent. Now you need to tune the prompt. Swap the model. Restrict access for specific users. Figure out what it's actually costing you.

README

MATE — The Command Center for AI Agents

Stop the redeploy-to-tweak loop.

Created by Ivan Antonijević

You built an agent. Now you need to tune the prompt. Swap the model. Restrict access for specific users. Figure out what it’s actually costing you. And make sure yesterday’s behavior still works after today’s changes.

Without a control layer, every one of those is a code change, a commit, and a redeploy.

MATE is that control layer. It adds everything production needs — live configuration, RBAC, cost tracking, regression testing, and an embeddable chat widget — without touching your agent code. Runs on Google ADK or LangGraph, switchable with one env var — same agents, same dashboard, same widget either way.


The Problem with Agents in Production

A single agent in a notebook is easy. A hierarchy of agents serving real users is not.

  • Iteration is slow. Tweaking a prompt means editing code, committing, and redeploying. By the time you’ve tested three variations you’ve wasted half a day.
  • Governance is messy. Who can call the admin agent? The finance agent? You need per-agent access control, but ADK doesn’t ship with RBAC.
  • Costs are invisible. You know tokens are being spent, but you can’t see which agent is the expensive one, or whether it’s prompt tokens, response tokens, or reasoning tokens eating your budget.
  • Regressions are silent. You improve one prompt and unknowingly break another agent’s logic. You only find out when a user complains.

Three Things MATE Gives You

🎨 The Studio — Build and iterate without touching code

A drag-and-drop canvas where you create agents, draw parent→child connections, and attach tools without writing JSON or Python. Change a prompt, swap a model (from Gemini to GPT-4o to local servers like Ollama, LM Studio, llama.cpp, LocalAI, or Llamafile), or restructure an entire hierarchy — and it takes effect immediately, no redeploy.

Every agent lives in the database. Every change is versioned. Roll back to any previous configuration in one click.

🏛️ The Control Room — Govern what runs in production

RBAC per agent means the finance team’s agent is never visible to basic users, without any custom middleware. Token tracking breaks down every request into four buckets — prompt, response, thoughts, and tool-use tokens — so you can see exactly where costs come from, per agent, per hour, per user.

Add guardrails (hallucination scoring, rate limits per user/project) from the dashboard. Audit logs capture every configuration change for compliance.

🧪 The Lab — Know before you ship

The Eval Framework runs your agents against a test suite automatically. Write expected outputs; MATE invokes the live agent, scores the response (exact match, semantic similarity, or LLM-as-Judge), and tracks pass rates across versions as a chart.

Set a regression threshold and MATE fires a webhook if a new version scores more than 5 points below the previous one. Catch prompt regressions before users do.


Quick Start

git clone https://github.com/antiv/mate.git && cd mate
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt

cp .env.example .env
# Set GOOGLE_API_KEY (or any supported provider key)

python auth_server.py
# Dashboard: http://localhost:8000  — login: admin / mate

Or with Docker:

docker-compose up

Migrations run automatically on startup. Default database is SQLite.


Screenshots

Dashboard Overview

Monitor usage, active agents, system health, and request trends at a glance.

Agent Management

View and manage your full agent hierarchy — models, parent relationships, and status per project.

Agent Visual Builder

Drag-and-drop canvas for building agent hierarchies. See tool and MCP nodes attached to each agent, click to configure them inline, and create connections without touching JSON.

Work Room

A built-in chat interface inside the dashboard. Pick any root agent from the card grid and start a conversation — no embed code, no browser tab switching. Sessions are auto-titled and persist across page reloads. The default landing page after login.

When the agent responds with code, the canvas panel opens automatically to the right of the chat. Code never clutters the conversation — it goes straight into a full-featured editor. HTML, JavaScript, CSS, and SVG can be executed in a sandboxed iframe with one click. Python runs via Pyodide (WebAssembly) directly in the browser. Dart and Flutter applications are executed directly in the browser via a seamless, zero-install DartPad integration that bypasses the code view to show the interactive application immediately. The panel is resizable by dragging the divider. Any edits made in the canvas are automatically included in the next prompt, so you can ask the agent to modify its own output without copy-pasting.

With the Canvas Panel, you can view and edit the generated code in a full-featured Ace Editor, and run/execute it in a sandboxed environment with a single click:

Editing generated code in the Canvas Panel

Previewing and running the generated code in a sandboxed iframe

Agent Configuration

Edit every aspect of an agent: model, instruction, RBAC roles, memory blocks, tools, MCP servers, planner, and schemas.

Tool Configuration

Toggle built-in tools or provide custom JSON — Google Drive, Search, Image, Memory Blocks, Code Executor, File Search, and more.

Chat Interface

Multi-agent chat with event tracing, tool calls, and persistent memory.

Usage Analytics

Token usage trends, agent performance, activity-by-hour, and per-agent success rates.

Token Usage Details

Drill into individual request logs — prompt / response / thought / tool-use token breakdown per request.


Why MATE vs. raw ADK

Challenge Raw ADK MATE
Change a prompt or model Edit code, redeploy Dashboard edit, instant
Switch LLM providers Code change per agent Change model_name in config (ollama_chat/llama3.2, lm_studio/qwen2.5, openai/gpt-4o, …)
Access control per agent Build your own Built-in RBAC, no code
Token cost visibility DIY 4-type tracking + analytics
Regression testing Manual Automated eval suite with LLM-as-Judge
Multi-team isolation Manual Project-scoped agent hierarchies
Embed chat on a website Not included Single `

A floating chat button appears instantly. No build step, no framework dependency.

### Multi-site example

MATE instance ├── key: wk_aaa → support_agent → company-support.com (English, light theme) ├── key: wk_bbb → sales_agent → company-sales.com (dark theme, no attachments) └── key: wk_ccc → docs_agent → docs.company.com (RAG over product docs)


Each key is independent: different agent, different greeting, different allowed origins, different color scheme.

### Page context awareness

When a user opens the widget on a product page, the widget automatically reads the page URL, title, and meta description and passes them to the agent as conversation context. The agent knows which page the user is viewing without any custom integration — just enable "Inject page context" in the widget admin panel.

### Widget admin panel

Each key has a standalone admin panel at `/widget/admin?key=...`. Non-technical teams can manage appearance, greeting, memory blocks, and uploaded files without touching the main dashboard.

| Setting | Admin panel |
|---|---|
| Widget title and greeting | ✓ |
| Light / dark / auto theme | ✓ |
| Button and accent color | ✓ |
| Show / hide file attachments | ✓ |
| Page context injection toggle | ✓ |
| Agent instruction and model | ✓ |
| Memory blocks (company info, FAQs…) | ✓ |
| RAG file upload | ✓ |

> Full documentation: [documents/WIDGET_INTEGRATION.md](documents/WIDGET_INTEGRATION.md)

---

## MCP Integration

MATE exposes agents and tools as MCP servers, compatible with Claude Desktop, Cursor, and any MCP client.

```bash
# Expose specific agents as MCP servers
export MCP_EXPOSED_AGENTS=creative_agent,support_agent

Each exposed agent gets endpoints at /agents/{name}/mcp/*. Built-in MCP servers: Image Generation (/images/mcp) and Google Drive (/gdrive/mcp).

Full documentation: documents/MCP_SERVERS.md


Dashboard Routes

Route Purpose
/dashboard/workroom Work Room — chat with any agent directly in the dashboard (default after login)
/dashboard Overview — usage stats, system health
/dashboard/agents Agent hierarchy, configuration, Visual Builder
/dashboard/users User management and role assignment
/dashboard/usage Token analytics, cost breakdown
/dashboard/evals Test suites, score history, regression tracking
/dashboard/audit-logs Audit trail viewer
/dashboard/migrations Database migration management
/dashboard/docs API documentation (Swagger + ReDoc)

Additional Documentation


Contributing

Contributions are welcome. See CONTRIBUTING.md for development setup, code style, and PR guidelines.

Security

For security best practices and vulnerability reporting, see SECURITY.md.

License

Apache License 2.0 — see LICENSE.

View this README on GitHub

インストール

This server does not publish a one-line install command.

Open the repository installation guide