Macha
Macha Concepts

Agent Evaluation

Definição

Agent Evaluation is Macha's built-in way to score how well an agent actually did its job. An AI judge runs a checklist over the agent's past conversations and returns, per conversation, whether it followed its instructions and why.

Também conhecido como: AI evaluationLLM-as-judgeinstruction-following evaluation

How it works

You set up a checklist once — or let Macha build one from the agent's own instructions — and an AI judge runs it over every past conversation in scope: chat sessions, autonomous trigger runs, sub-agent handoffs, embedded chatbot conversations. Two fields are always present: Instructions followed (Yes / Partially / No) and a short Why. You add your own checks, like "right tool called" or "refund granted only when policy allowed."

The headline score gives partial credit — Yes count plus half the Partially count, over the total — and results come back as a table, a report with charts, and a per-conversation reason so you can audit a verdict without re-reading the whole thread.

Auto-Eval vs manual evaluations

Every workspace gets Auto-Eval free on every plan: from the moment you create an agent, Macha grades its conversations against its live instructions on a hardcoded judge, up to 50 conversations per agent, at no credit cost. Manual evaluations (Professional and Enterprise) are the configurable flow — pick the scope, define a custom rubric, choose the judge model, and evaluate against a specific version of the instructions — billed one credit per evaluated conversation at the judge model's rate.

This is how you measure a prompt change: run before and after and compare the scores, instead of guessing whether the edit helped. It evaluates past conversations, not in-flight ones — for real-time control, use the agent's instructions and confirmation gates.

Perguntas frequentes

What does the judge actually check?

Whether the agent followed its own instructions on each past conversation — the right tools, the right routing, the tone and boundaries you set — plus any custom checks you add. It returns a Yes/Partially/No verdict and a short reason per conversation.

Do I have to pay to evaluate my agents?

Auto-Eval is free on every plan (up to 50 conversations per agent). Manual, configurable evaluations are a Professional and Enterprise feature and cost one credit per evaluated conversation.

Coloque essas ideias em prática

A Macha é uma camada de agentes de IA que fica em cima do help desk que você já usa, Zendesk, Freshdesk, Front, Intercom ou Gorgias.

Começar teste