Macha
Agent Simulations

Run the agent on real tickets. Let it change nothing.

A simulation puts your agent through your actual production tickets and shows you exactly what it would have done on each one — without sending a single reply or firing a single write.

The real ticket production

#20831 · Where is my order?

Antonio R. · open · 2 replies

I ordered last week and still haven’t received anything. Order number 20831. Can you help?

Unchanged. No reply, no tag, no status change.
What the agent would have done Instructions followed
read Looked up order 20831 in Shopify — shipped 12 Jul, out for delivery.
read Searched the help centre for the late-delivery policy.
held Would have posted this reply and tagged the ticket “shipping”. Not sent.

The reply it would have sent

“Hi Antonio — your order shipped on 12 July and is out for delivery today, arriving by 6pm. If it hasn’t landed by then, reply here and I’ll chase the courier straight away.”

📦

Refund Agent

@refund · GPT-5

Active
Instructions written
6 tools wired
Trigger configured

Never run on a real ticket

You built the agent. Now you have to trust it with a real customer.

That is the moment everyone stalls. The instructions read well, it behaved in the test panel, and the sample tickets you tried were the ones you happened to think of. None of that tells you how it handles the messy, half-explained, three-topics-in-one tickets that actually arrive.

So the agent sits switched off, or it goes live and someone watches the queue with their finger over the pause button. Simulations exist to remove that choice: run it on the real thing first, and read what it would have done before anything reaches anyone.

How it stays safe

Reads are real. Writes are captured, never sent.

A dry run is only useful if the context is genuine, so the agent really does query your data. The difference is what happens when it tries to act: every write is intercepted and answered with a plausible response, so the agent carries on reasoning while nothing lands.

Read tools run for real

Live data · no side effects

The agent sees exactly what it would see in production, which is the only way the run tells you anything true.

Searches your knowledge base
Reads the ticket, its fields and the whole thread
Fetches live order and account data from connected tools

Write tools are intercepted

Zero side effects

Replies, status changes and assignments are captured so you can read them, then thrown away. The agent believes the write succeeded and keeps reasoning; nothing leaves your workspace.

Posting a reply to the ticket
Changing status, priority or assignee
Adding tags, notes or triggering a refund

Which means you can point a brand-new agent at last week’s real queue on day one, and the worst thing that happens is you learn something.

The part that makes it honest

Every ticket is replayed as if it had just come in.

A solved ticket already contains the answer, so replaying it whole would prove nothing. Macha strips the ticket back to its first customer message and hands the agent that — with the original subject, tags and priority, and nothing else.

What the agent sees

The customer’s first message
The original subject
Tags and priority as they were

What is stripped out

Every later reply in the thread
Agent notes and internal comments
Status changes and the resolution

So the run tells you how your agent handles a fresh incoming ticket — not how well it can read someone else’s answer.

When it earns its keep.

Any time the honest answer to “will this work?” is “probably, but I would rather not find out in front of a customer”.

A new agent, before it ever goes live

You have written the instructions and wired the tools. Run it across a few hundred real tickets and read the failures before the trigger is ever switched on.

A new customer’s real ticket mix

Prove the agent handles their queue, not synthetic examples, without touching a single one of their live conversations.

A tool wiring you are unsure about

Every call the agent makes is listed with its arguments and the response it got. Missing tools and wrong parameters surface here instead of in production.

What you get back.

The same judge that grades your live conversations grades every simulated one, so the numbers are directly comparable to the agent’s real adherence score.

Adherence

92%

Records

120

Writes captured

206

Credits

480

Yes

#20831 · Where is my order?

Looked up the order, gave the tracking date, offered to chase the courier. Correctly did not offer a refund.

Yes

#20847 · Wrong size shipped

Confirmed the item, started a return, would have tagged the ticket and set it pending.

Partially

#20852 · Discount code failed

Answered correctly but skipped the instruction to check whether the code had already been redeemed.

Simulated conversations are stored separately — they never count towards your live adherence and never appear in your team’s past conversations.

Questions about Simulations.

Short answers. If something is missing, ask us — we will tell you straight.

No. Every write — posting a reply, changing status, tagging, refunding — is intercepted by the tool executor and answered with a canned response. Only read tools reach your data, so the run has real context and leaves no trace.
Because you choose the test tickets there, and you choose the ones you have thought about. A simulation takes whatever actually arrived in a date range, including the confusing ones, and runs the agent on all of them.
The results stand. The run freezes the model at launch and snapshots what it used, so a later prompt or model change does not quietly invalidate what you already measured.
No. They are stored with their own source and filtered out of the evaluation view, so your headline number reflects real traffic only.
Yes — cap the run at a small number of records. You see the credit estimate before launching, and optional skip filters drop auto-replies before you spend anything on them.

Bereit, dein Team mit KI zu turbo-boosten?

Starte in Minuten. Verbinde deine Tools, konfiguriere deine Agenten und lass die KI den Rest erledigen.

500 Freikredite · ohne Zeitlimit, ohne Kreditkarte