Macha

Choosing the Right Model for Your Agent: GPT-5, GPT-5.4, Mini & Credits

Abbas, Customer Support & AI, Macha

Written by

Ankeet Guha, Co-founder & CTO, Macha

Reviewed by

Published July 26, 2026

Updated July 26, 2026

Every Macha agent runs on a model you pick, and that single choice quietly sets two things at once: how good the agent's answers are, and how much each answer costs you. Pick too small a model for a delicate billing-dispute agent and it fumbles the edge cases. Pick the flagship for a high-volume "where is my order" bot and you burn three times the credits for a reply nobody could tell apart. Most teams never revisit the default — and leave both quality and budget on the table.

Choosing the Right Model for Your Agent: GPT-5, GPT-5.4, Mini & Credits

This is a decision guide, not a spec sheet. Macha gives every agent its own model setting (and every sub-agent its own on top of that), so the right move is rarely "use the best model everywhere." It's matching each agent to the smallest model that clears the job, and spending the saved credits where the work is genuinely hard. Below is how the current lineup actually differs, what each tier costs per response, and a simple framework for choosing — including when not to reach for the big model.

First, the four levers (this part isn't Macha-specific)

Strip away the brand names and model selection is the same trade-off on any platform that lets you choose a model per task — Macha, a raw OpenAI or Anthropic API, or anything in between. Four levers pull against each other, and picking a model is just deciding where you want to sit on each:

  • Quality / reasoning — how well the model handles nuance, ambiguity, and strict multi-step logic. The thing you're tempted to max out everywhere, and the thing you rarely need to.
  • Cost — here it's credits per response; on a raw API it's tokens; on a per-seat tool it's licences. A flagship can cost two to three times the small model for a reply a customer couldn't tell apart.
  • Latency — how fast the answer comes back. Smaller models are faster, which matters most for live chat and anything high-volume where a queue can back up.
  • Context window — how much the model can hold at once: a long ticket thread, several knowledge sources, a big system prompt. Bigger windows stop a model losing the plot on messy threads, but you pay for the headroom.

The whole game is spending quality and context only where a mistake is expensive, and buying cheap speed everywhere else. Macha's contribution is making that choice explicit and per-agent instead of a single global default. With that lens, here's the actual lineup.

The lineup, and what a credit buys

Macha bills in credits, and a credit is charged per complete assistant response — not per word, not per ticket, and not per "resolution." One response from an agent costs whatever its model costs, deducted once when the reply finishes. That makes model choice the single biggest lever on your monthly credit burn.

Here's the current lineup and the credit cost per response. The OpenAI tiers got cheaper after the June 17 price cut that dropped the mid and top tiers — and Macha also exposes Claude and Groq models in the same picker:

ModelProviderCredits / responseContextBest for
GPT-5.4 MiniOpenAI1 (default)400KHigh-volume agents: WISMO, status, triage, classification, FAQ deflection
GPT-5OpenAI21MThe balanced default for nuanced replies at volume — most general-purpose agents
GPT-5.4OpenAI31MStrict, high-stakes logic: complex branching, "do this exactly once" rules, refunds/billing
Claude Sonnet 4.5Anthropicsee /pricing†LargeWarm, emotionally-aware replies on sensitive tickets; image vision
Claude Sonnet 4Anthropicsee /pricing†LargeThe same calm tone with vision, a step down from 4.5 on cost
Groq (Llama-class)Groqsee /pricing†SmallerUltra-low latency, cheapest per response — but no image vision
†Across the whole platform, per-response cost ranges from roughly 0.5 to 9 credits by model; the Claude and Groq models sit inside that band. The exact figure for each is shown live on the model picker and on the pricing page.

A few things worth knowing before you choose:

  • Mini is the default for a reason. New agents start on GPT-5.4 Mini at 1 credit. It has a 400K-token context window, supports image vision, and follows instructions well enough for the large majority of support workflows. Most of your fleet should probably stay here.
  • GPT-5 got cheaper. It dropped from 3 credits to 2 on June 17, which changes the math: the "balanced" model is now only double the cost of Mini, not triple. It's the natural step up when Mini's answers feel a little thin but you don't need flagship reasoning.
  • GPT-5.4 is the instruction-follower. It dropped from 5 credits to 3. Its edge over GPT-5 isn't raw eloquence — it's discipline: noticeably better at honoring strict rules ("only ever issue one refund," "never promise a delivery date"), complex branching, and multi-step tool sequences where order matters.
  • Claude is the tone pick. The two Claude Sonnet models (4.5 and 4) are vision-capable and, per industry consensus, hold a calmer, more human tone on emotionally charged tickets (gettalkative) — reach for them on refund-gone-wrong or complaint threads where how the reply reads matters as much as what it says.
  • Groq is the speed-and-cost pick. Groq's Llama-class models are the fastest and cheapest in the picker, which suits very high-volume, latency-sensitive queues — with one hard limit: they can't read image attachments and return a "not supported" message, so don't put a vision agent on them.
A note on retired models: if you've been on Macha a while, you may remember GPT-5 Mini, GPT-4o Mini, and o4-mini. Those were retired on April 23 and every agent on them was auto-migrated to GPT-5.4 Mini — so there's nothing to fix, but it's worth re-checking those agents to confirm Mini is still the right fit.

Where you set the model (and the gotcha)

There are two places a model gets chosen, and confusing them is the most common mistake we see.

The agent's default lives in the agent's configuration card — model selection moved here in the April builder refresh, with a helper notice spelling out the differences. This is the model the agent uses every time it runs autonomously from a trigger (a new Zendesk ticket, a scheduled job). This is the one that matters for cost at scale.

A Macha agent's configuration where you set its default model alongside instructions, tools, and triggers.
A Macha agent's configuration where you set its default model alongside instructions, tools, and triggers.

The per-conversation override is the dropdown in the top-right of a chat. Change it there and you've changed the model for that one conversation only — Macha now pops a toast making this explicit and pointing you to Settings if you actually meant to change the org default. Handy for spot-testing whether a tougher model handles a tricky thread better before you commit the agent to it.

Because the agents table now carries a Model column with each provider's icon, you can audit your whole fleet at a glance — and that's the fastest way to spot the expensive mistake: a low-value, high-volume agent quietly sitting on the flagship.

The Macha agents list with a Model column showing each agent's provider and tier at a glance.
The Macha agents list with a Model column showing each agent's provider and tier at a glance.

A framework: match the model to the job

Forget "which model is best." Ask "how expensive is a mistake here, and how much volume runs through this agent?" That two-axis read gets you to the right tier almost every time.

Start on Mini (1 credit) when…

The task is bounded and a wrong answer is cheap to recover from:

  • WISMO / order status — look up an order, report the tracking. (See the order-status use case →)
  • Triage and routing — read a ticket, pick a category, set a tag, assign a group.
  • Classification and tagging — turn free-text into structured fields.
  • FAQ deflection — answer from a knowledge source you've grounded it on.

These are high-volume and forgiving. At Mini's 1 credit, a Professional plan's 10,000 credits covers ten thousand of these responses a month. There is rarely a reason to pay more.

Step up to GPT-5 (2 credits) when…

The agent writes customer-facing prose that has to read well, reasons over a longer or messier thread, or stitches a few tools together:

  • A front-line resolution agent that drafts the actual reply a customer reads.
  • Anything juggling a long conversation history or several knowledge sources where Mini starts to lose the thread.
  • General-purpose agents where you've noticed Mini's answers are fine but not good.

At 2 credits you get materially better reasoning for double the cost — and after the June price cut that's a smaller jump than it used to be.

Reserve GPT-5.4 (3 credits) for…

The agents where a single mistake is genuinely expensive and the logic is strict:

  • Refunds, billing, and policy agents bound by hard rules ("refund once, never twice," "escalate anything over $X").
  • Complex branching workflows with many conditional steps and tool calls that must fire in the right order.
  • Multi-step orchestration where an off-by-one error compounds.

This mirrors the wider industry pattern: route the bulk of traffic to a small fast model and escalate only the hard 10–20% to a flagship, going premium only when quality demonstrably suffers on the cheaper tier (aicomparison.ai, Cobbai). Macha lets you implement exactly that — per agent, and per sub-agent.

The sub-agent trick

Multi-agent setups make this concrete. A cheap Mini router reads every incoming ticket and delegates: most go to Mini specialists, but a billing dispute or a multi-part complaint gets handed to a GPT-5.4 sub-agent. You pay flagship credits only on the fraction of tickets that need them, while the front door stays at 1 credit. Each sub-agent carries its own model, so this routing is just configuration, not code.

What it costs: three worked plans

To make the trade-off tangible, here's how far a Professional plan's 10,000 credits stretches at each tier, assuming an agent that produces one response per ticket:

ModelCredits / responseResponses in a Pro plan's monthly credits
GPT-5.4 Mini1~10,000
GPT-52~5,000
GPT-5.43~3,333

The headline: moving a high-volume agent from the flagship down to Mini triples its response capacity for the same spend — with no perceptible quality loss on bounded tasks. Conversely, paying for GPT-5.4 on a refund agent that handles a few dozen sensitive tickets a day costs almost nothing in the scheme of things and buys you real safety. You can watch the burn live on the Billing page's usage card, which now also shows any never-expiring top-up credits stacked above your monthly balance.

The Macha billing page usage card showing monthly credit consumption against the plan allowance.
The Macha billing page usage card showing monthly credit consumption against the plan allowance.

Don't pick on vibes — measure it

The honest answer to "is Mini good enough here?" is test it on your real tickets. Two ways to do that inside Macha:

  1. Per-conversation override. Open a representative thread, switch the model in the top-right dropdown, and re-run the agent. Compare the two replies side by side. Cheap, instant, no commitment.
  2. A Studies or eval run. For a rigorous read, run the same prompt across a batch of historical tickets on Mini, then on GPT-5, and compare the outputs at scale before you change a production agent. (See our walkthrough of running an AI analysis across thousands of tickets.)

Both beat guessing. A model that looks great on three test cases can drift under volume — production consistency and latency matter as much as peak quality, which is exactly why the support-LLM literature pushes evaluation over intuition (eesel AI).

Watch-outs: when not to upgrade

A guide that only ever says "use the bigger model" is selling you credits. Some honest caveats:

  • Bigger isn't always better-behaved. GPT-5.4's strength is instruction discipline, not prose. If your agent's job is a warm, human-sounding apology, a Claude model or even GPT-5 may feel better to customers than the flagship. Match the model to the failure mode you actually have.
  • The flagship won't fix a bad prompt. If an agent hallucinates because it has no grounding, upgrading the model wastes credits — fix the knowledge source and instructions first. Model choice is the last tuning knob, not the first.
  • Vision narrows your options. If an agent must read image attachments on Zendesk tickets, you're limited to vision-capable models (GPT-5.4 Mini and the Claude Sonnet models); Groq models return a "not supported" message. Check capability before cost.
  • Defaults drift. Auto-migrated agents and copied-from-template agents can quietly sit on a model that no longer fits. Audit the agents table's Model column quarterly.
  • Per-conversation ≠ org default. Changing the model in a chat does not change what your trigger-fired agent uses in production. Set the default in the agent's config; the chat dropdown is for testing.

FAQ

What's the default model for a new Macha agent? GPT-5.4 Mini, at 1 credit per response. It's fast, supports image vision, has a 400K-token context window, and handles most support workflows — so most agents should stay on it unless you have a reason to move up.

How are credits charged — per ticket or per message? Per complete assistant response. A credit (or two, or three, by model) is deducted once when the agent finishes a reply. Enterprise plans bypass credit checks entirely. Credits measure AI actions, not deflections or resolutions.

What's the real difference between GPT-5 and GPT-5.4? GPT-5 (2 credits) is the balanced all-rounder for nuanced, customer-facing replies at volume. GPT-5.4 (3 credits) is the highest-capability OpenAI model in Macha, with notably better instruction-following — pick it for strict rules, complex branching, and high-stakes logic, not just for "better writing."

Can different agents use different models? Yes — every agent has its own model setting, and every sub-agent has its own on top of that. The common pattern is a cheap Mini router that escalates only hard tickets to a flagship sub-agent.

Does changing the model in chat change it everywhere? No. The top-right dropdown changes the model for that one conversation only. To change what an agent uses in production, set its default model in the agent's configuration (Macha shows a toast reminding you of this).

Which models can read images? GPT-5.4 Mini and the Claude Sonnet models (4.5 / 4) can analyze image attachments on Zendesk tickets. Groq models gracefully decline. If image vision matters for an agent, choose from the vision-capable set.

Pick once, then audit

Model choice isn't a set-and-forget decision — it's a dial you tune per agent against two questions: how costly is a mistake, and how much volume flows through. Default to Mini, step up to GPT-5 where the writing or reasoning needs it, and reserve GPT-5.4 for the strict, high-stakes agents. Then use the agents table to make sure nothing's quietly overspending.

Want to try it on your own queue? Start a 7-day free trial, no credit card required, connect your helpdesk, and switch a couple of agents between models to feel the difference — or read the docs and pricing for the full credit breakdown.


Written by Abbas (Customer Support & AI, Macha) · Reviewed by Ankeet Guha (Co-founder & CTO) · Published 2026-06-24 · Last updated 2026-06-24.

Macha

About Macha

Macha is an AI agent platform that works on top of the help desk you already use — Zendesk, Freshdesk, Gorgias, or Front — and connects to the rest of your stack, even your own internal systems. Its AI agents resolve tickets and automate entire workflows end to end, all set up in plain English, no code. Learn more about Macha →

Zendesk
5.0 on Zendesk Marketplace

Loved by support teams worldwide

See what support teams are saying about Macha AI.

The application seems excellent to me! We are still testing, and we need support for some details and they were extremely efficient too!

Daniela Costa

Daniela Costa

Head of Support, Seabra

Macha has been a great addition to our support toolkit. It generates clear, well-organized responses that fit naturally into our workflow. One feature we particularly appreciate is its ability to automatically reply in the same language as the ticket.

Marius F

Marius F

Support Head, Zentana

We've been using Macha for a little while now and it's been really great addition so far! It's powerful, convenient, and makes getting work done a lot easier for our agents.

Alexander Wedén

Alexander Wedén

Head of Support

Support team is very helpful and responsive. Really enjoy how lightweight this is within Zendesk itself vs other more intrusive tools.

Cathleen Wright

Cathleen Wright

Zendesk Admin, Cortex IO

So far it's pretty good! Our queries are a little nuanced, so we can't always use it, but it's got enough utility for us. It can even incorporate our bilingual country with greetings in a second language.

Jae Oliver

Jae Oliver

Head of Support, Wise

Really enjoying using Macha, it has made a noticeable difference to our support team in a short amount of time. I really like the ticket summary feature, saves us a lot of time.

Harry Jackson

Harry Jackson

Head of Support, Crumb

Macha AI is a great addition to my workspace! It's powerful, convenient, and it really makes productivity so much easier for our agents!

Dave G

Dave G

Head of Support, Cyber Power Systems

Very impressed! AI integration for Zendesk has certainly come a long way and Macha seems to set the standard for now. This will for sure save lot of time in our support team.

Pauli Juel

Pauli Juel

Head of CS, Dokument24

Macha has been working great for us so far! The auto-responses are accurate and our resolution time has dropped significantly.

Lana T

Lana T

Zendesk Admin, Swotzy

Macha AI is a great addition. The knowledge base feature means our agents always have the right answers at their fingertips.

Mischa Wolf

Mischa Wolf

Head of Support, Topi

We're enjoying this integration so far. It's made our support team more efficient and our customers get faster responses.

Paula G

Paula G

Head of Customer Support, Xly Studio

The team enjoys using it. It saves considerable time on common questions and the integration options are excellent.

Kilian Leister

Kilian Leister

Support Head, Didriksons

Ready to supercharge your team with AI?

Get started in minutes. Connect your tools, configure your agents, and let AI handle the rest.

500 free credits · no time limit, no credit card