Macha
Macha Concepts

Agent Evaluation

Definición

Agent Evaluation is Macha's built-in way to score how well an agent actually did its job. An AI judge runs a checklist over the agent's past conversations and returns, per conversation, whether it followed its instructions and why.

También conocido como: AI evaluationLLM-as-judgeinstruction-following evaluation

How it works

You set up a checklist once — or let Macha build one from the agent's own instructions — and an AI judge runs it over every past conversation in scope: chat sessions, autonomous trigger runs, sub-agent handoffs, embedded chatbot conversations. Two fields are always present: Instructions followed (Yes / Partially / No) and a short Why. You add your own checks, like "right tool called" or "refund granted only when policy allowed."

The headline score gives partial credit — Yes count plus half the Partially count, over the total — and results come back as a table, a report with charts, and a per-conversation reason so you can audit a verdict without re-reading the whole thread.

Auto-Eval vs manual evaluations

Every workspace gets Auto-Eval free on every plan: from the moment you create an agent, Macha grades every autonomous run against its live instructions on a hardcoded judge, with no conversation cap and at no cost. Manual evaluations, also included on every plan, are the configurable flow — pick the scope, define a custom rubric, choose the judge model, and evaluate against a specific version of the instructions.

This is how you measure a prompt change: run before and after and compare the scores, instead of guessing whether the edit helped. It evaluates past conversations, not in-flight ones — for real-time control, use the agent's instructions and confirmation gates.

Preguntas frecuentes

What does the judge actually check?

Whether the agent followed its own instructions on each past conversation — the right tools, the right routing, the tone and boundaries you set — plus any custom checks you add. It returns a Yes/Partially/No verdict and a short reason per conversation.

Do I have to pay to evaluate my agents?

No. Auto-Eval is free on every plan and grades every autonomous run, with no conversation cap. Manual, configurable evaluations are included on every plan too.

Pon estas ideas a trabajar

Macha es una capa de agentes de IA que va sobre el help desk que ya usas, Zendesk, Freshdesk, Front, Intercom o Gorgias.

Empezar prueba