Run the agent on real tickets.
Let it change nothing.
A simulation puts your agent through your actual production tickets and shows you exactly what it would have done on each one — without sending a single reply or firing a single write.
#20831 · Where is my order?
Antonio R. · open · 2 replies
I ordered last week and still haven’t received anything. Order number 20831. Can you help?
The reply it would have sent
“Hi Antonio — your order shipped on 12 July and is out for delivery today, arriving by 6pm. If it hasn’t landed by then, reply here and I’ll chase the courier straight away.”
Refund Agent
@refund · GPT-5
Never run on a real ticket
You built the agent. Now you have to trust it with a real customer.
That is the moment everyone stalls. The instructions read well, it behaved in the test panel, and the sample tickets you tried were the ones you happened to think of. None of that tells you how it handles the messy, half-explained, three-topics-in-one tickets that actually arrive.
So the agent sits switched off, or it goes live and someone watches the queue with their finger over the pause button. Simulations exist to remove that choice: run it on the real thing first, and read what it would have done before anything reaches anyone.
How it stays safe
Reads are real. Writes are captured, never sent.
A dry run is only useful if the context is genuine, so the agent really does query your data. The difference is what happens when it tries to act: every write is intercepted and answered with a plausible response, so the agent carries on reasoning while nothing lands.
Read tools run for real
Live data · no side effects
The agent sees exactly what it would see in production, which is the only way the run tells you anything true.
Write tools are intercepted
Zero side effects
Replies, status changes and assignments are captured so you can read them, then thrown away. The agent believes the write succeeded and keeps reasoning; nothing leaves your workspace.
Which means you can point a brand-new agent at last week’s real queue on day one, and the worst thing that happens is you learn something.
The part that makes it honest
Every ticket is replayed as if it had just come in.
A solved ticket already contains the answer, so replaying it whole would prove nothing. Macha strips the ticket back to its first customer message and hands the agent that — with the original subject, tags and priority, and nothing else.
What the agent sees
What is stripped out
So the run tells you how your agent handles a fresh incoming ticket — not how well it can read someone else’s answer.
When it earns its keep.
Any time the honest answer to “will this work?” is “probably, but I would rather not find out in front of a customer”.
A new agent, before it ever goes live
You have written the instructions and wired the tools. Run it across a few hundred real tickets and read the failures before the trigger is ever switched on.
A new customer’s real ticket mix
Prove the agent handles their queue, not synthetic examples, without touching a single one of their live conversations.
A tool wiring you are unsure about
Every call the agent makes is listed with its arguments and the response it got. Missing tools and wrong parameters surface here instead of in production.
What you get back.
The same judge that grades your live conversations grades every simulated one, so the numbers are directly comparable to the agent’s real adherence score.
Adherence
92%
Records
120
Writes captured
206
Credits
480
#20831 · Where is my order?
Looked up the order, gave the tracking date, offered to chase the courier. Correctly did not offer a refund.
#20847 · Wrong size shipped
Confirmed the item, started a return, would have tagged the ticket and set it pending.
#20852 · Discount code failed
Answered correctly but skipped the instruction to check whether the code had already been redeemed.
Four features that sound alike. They are not.
A Simulation runs your agent on real records and shows what it would have done. Nothing reaches the customer.
Questions about Simulations.
Short answers. If something is missing, ask us — we will tell you straight.
Pronto para turbinar seu time com IA?
Comece em minutos. Conecte suas ferramentas, configure seus agentes e deixe a IA cuidar do resto.
500 créditos grátis · sem limite de tempo, sem cartão de crédito
Shopify
Stripe
Slack
Notion
Google Workspace
Confluence