Macha
Continuous evaluation

Grade every conversation, not just the obvious ones.

Every conversation your agent handles is graded against its own instructions. See exactly which rule was followed and which one slipped, the moment it slips.

Grade for conv_6a5a1206 · @refundAgent just now
Instructions followed Yes

The agent pulled the ticket, checked the customer's order history, and matched the refund window before replying.

Followed

Always confirm the customer's order details before quoting policy.

It offered a refund without asking the reason for return, which the instructions require.

Not followed

Before approving a refund, ask the customer for the reason for return.

Ticket 8421 · refund reply

Yes · followed refund policy check

2m

Ticket 8420 · shipping question

Partially · skipped tracking link

4m

Ticket 8419 · return request

Yes · asked reason for return

7m

Ticket 8418 · escalation

No · missed handoff step

11m

Grade every conversation.

  • Yes / Partially / No

    A clear verdict on every conversation with a written rationale.

  • Free with every plan

    Continuous evaluation runs on every conversation, at no extra cost to you.

  • Graded against the right version

    Judged against the instructions that were live when the conversation happened, not what you've since edited.

The judge quotes the exact rule.

  • Verbatim citations

    When a rule is cited, it's quoted directly from your instructions. No paraphrasing.

  • Followed, not-followed, ambiguous

    Every rule the judge references is tagged so you can see at a glance what stuck and what slipped.

  • Read it like a review, not a rating

    Rule + reasoning together, so a reviewer knows exactly why a call was right or wrong.

Why · @wismoAgent · ticket 8421

Read the ticket, then went straight to the Shopify order lookup as instructed.

Followed

For every shipping question, look up the order before replying.

Included the tracking link and the delivery window in the same message.

Followed

Share the tracking URL and the estimated delivery window in the same reply.

The one slip: left the ticket status open when a pending status would have been more accurate.

Not followed

If the customer is still awaiting delivery, set the ticket to pending before replying.

Watch adherence over time.

Every graded conversation contributes to a rolling adherence score. When it moves, you know exactly which rule started slipping.

Adherence score

86%

↑ 4% vs last week
Yes 312
Partially 48
No 11

Last 30 days

@refundAgent
371 conversations judged
Jun 18Jun 25Jul 2Jul 9Today

Top failing rule · 22 conversations

Before approving a refund, ask the customer for the reason for return.

Started slipping on July 2. Reverting the model on the sample set restored it within a day.

Built with Claude Code

Extend it with a coding agent.

Continuous evaluation is a live signal you can wire into whatever comes after. A nightly summary. A Slack ping when adherence drops. A weekly report for your team.

Hand your Macha API key to Claude Code or Codex and it can read the shape of your evaluations and build the workflow you have in mind, without you writing the plumbing.

Build with Claude Code & Codex
Slack

#support-quality

Macha evals bot · 9:00 AM

Weekly adherence for @refundAgent

Yes

312

Partially

48

No

11

Top slip: "ask reason for return", 22 conversations.

How every grade lands

Conversation

@refundAgent

Called get_ticket, checked order history, sent reply.

1

Conversation completes

Chat, autonomous, embed, or hand-off. The moment the run finishes, the grade fires in the background.

Y
Followed
N
Not followed
GPT-5 judge
2

Judge grades it

The judge reads the full trace and quotes the exact rule that was followed and the exact one that slipped.

Rolling score
Top failing rule
Drift alerts
3

Watch adherence move

A rolling score, the top failing rule, and drift alerts. Everything queryable, everything exportable.

Questions about evaluation.

Everything you need to know about grading every conversation your agent handles.

Every conversation your agent handles is graded against its own instructions by an LLM judge. The judge picks a verdict (Yes, Partially, or No), quotes the exact rules it saw followed or slipped, and stores every judgment as a record you can query, group, and export.
No. Continuous evaluation is included free on every plan, including trials. The judgments do not consume any of your credits, so you can keep it on across every agent, every conversation, without watching the meter.
The judge reads the full trace of a completed conversation, then answers one question: did the agent follow its own instructions? It scores against the instructions that were live when the conversation happened, quotes the rule it is citing verbatim, and tags each cited rule as followed, not followed, or ambiguous.
The judge and its schema are locked so results are comparable across agents and over time. What you customize is the instructions on your agent itself. Every rule you write is a rule the judge will grade against on future conversations. If you want a per-team or per-workflow variant, you can also spin up your own manual studies alongside continuous evaluation.
No. The grade runs after the conversation completes, fire-and-forget, in the background. Your agent replies to the customer without waiting for the judge, and the judgment lands on its own once the run has finished.
Zendesk
5.0 on Zendesk Marketplace

Loved by support teams worldwide

See what support teams are saying about Macha AI.

The application seems excellent to me! We are still testing, and we need support for some details and they were extremely efficient too!

Daniela Costa

Daniela Costa

Head of Support, Seabra

Macha has been a great addition to our support toolkit. It generates clear, well-organized responses that fit naturally into our workflow. One feature we particularly appreciate is its ability to automatically reply in the same language as the ticket.

Marius F

Marius F

Support Head, Zentana

We've been using Macha for a little while now and it's been really great addition so far! It's powerful, convenient, and makes getting work done a lot easier for our agents.

Alexander Wedén

Alexander Wedén

Head of Support

Support team is very helpful and responsive. Really enjoy how lightweight this is within Zendesk itself vs other more intrusive tools.

Cathleen Wright

Cathleen Wright

Zendesk Admin, Cortex IO

So far it's pretty good! Our queries are a little nuanced, so we can't always use it, but it's got enough utility for us. It can even incorporate our bilingual country with greetings in a second language.

Jae Oliver

Jae Oliver

Head of Support, Wise

Really enjoying using Macha, it has made a noticeable difference to our support team in a short amount of time. I really like the ticket summary feature, saves us a lot of time.

Harry Jackson

Harry Jackson

Head of Support, Crumb

Macha AI is a great addition to my workspace! It's powerful, convenient, and it really makes productivity so much easier for our agents!

Dave G

Dave G

Head of Support, Cyber Power Systems

Very impressed! AI integration for Zendesk has certainly come a long way and Macha seems to set the standard for now. This will for sure save lot of time in our support team.

Pauli Juel

Pauli Juel

Head of CS, Dokument24

Macha has been working great for us so far! The auto-responses are accurate and our resolution time has dropped significantly.

Lana T

Lana T

Zendesk Admin, Swotzy

Macha AI is a great addition. The knowledge base feature means our agents always have the right answers at their fingertips.

Mischa Wolf

Mischa Wolf

Head of Support, Topi

We're enjoying this integration so far. It's made our support team more efficient and our customers get faster responses.

Paula G

Paula G

Head of Customer Support, Xly Studio

The team enjoys using it. It saves considerable time on common questions and the integration options are excellent.

Kilian Leister

Kilian Leister

Support Head, Didriksons

Ready to supercharge your team with AI?

Get started in minutes. Connect your tools, configure your agents, and let AI handle the rest.

500 free credits · no time limit, no credit card